The 30-second read
Grounding: tools on vs. tools off
Grounding is the share of question-specific facts that could only have come from the synthetic tool data. The saved sweep scored every tools-off answer at 0.00.
Bars show tools-on grounding. Tools-off grounding was 0.00 for all four questions.
| Question | Tools on | Tools off | Delta |
|---|---|---|---|
| drug_tier | 0.50 | 0.00 | +0.50 |
| drug_flags | 0.75 | 0.00 | +0.75 |
| drug_absent | 0.43 | 0.00 | +0.43 |
| ed_surge | 0.67 | 0.00 | +0.67 |
Native tool call: formulary lookup
Question
What is the formulary tier for zephadril, and does it require prior authorisation?
Tool selected
get_drug_info(name="zephadril")
Grounded answer
Zephadril is synthetic formulary tier 3 and does require prior authorisation.
Why the trace matters
The model emitted a native structured tool call, Python returned the synthetic JSON record, and the final answer cited that result. The evidence chain is visible rather than inferred from the answer.
Saved run: 2 model calls, 1,381 prompt tokens, 154 completion tokens.
Open the saved native trace
--------------------------------------------------------------------------
COMPARE · run_tool_trace (same question, same model, same cache)
agree path_used
agree stage sequence
agree tools used
agree produced an answer
A stage-sequence difference is usually one missing make_stage() call; a
path_used difference usually means you set it from what you INTENDED rather
than from what happened. Both matter: the trace is the evidence.
--------------------------------------------------------------------------
Steps 4-5 will use the INSTRUCTOR loop. To run them on your own, change the line
above to: RUN_TRACE = run_tool_trace
============================================================================
TRACE · drug_tier
tools ON · path used: native
model openrouter/free via openrouter · 2 model call(s) (2 from cache) · 0.0 s · 1381 in / 154 out tokens · $0.000000
============================================================================
[1] PROMPT (turn 0) the prompt you sent [native]
system : You are a hospital formulary assistant used in a graduate analytics class (MIS 752, Lab 6). You know NOTHING about the drugs or the arrival forecasts in this conversation except what your tools return. Rules: (1) if a tool can answer, call it instead of guessing; (2) never invent a drug fact, a formulary tier, a drug class or a number; (3) if a tool says there is no record, say plainly that there is no record; (4) answer in at most four sentences and name the tool result you are relying on.
user : What is the formulary tier for zephadril, and does it require prior authorisation?
[2] MODEL_TURN (turn 1) model response, turn 1 [0.92 s · 597 in / 69 out tok · native]
[3] TOOL_CALL (turn 1) the model asked for a tool [native]
name : get_drug_info
arguments : {"name": "zephadril"}
raw : {"name": "zephadril"}
[4] TOOL_RESULT (turn 1) your Python function answered [0.00 s · native]
your function returned, in 0.03 ms:
{"drug": "zephadril", "found": true, "synthetic": true, "warning": "SYNTHETIC TEACHING STUB -- these drug names are fictional and every value this stub returns was invented for MIS 752 Lab 6. Inside this lab the stub's answer is authoritative; outside it, nothing here is real. Never use for a clinical decision.", "source": "MIS 752 Lab 6 invented formulary v1", "drug_class": "invented class 7A (teaching placeholder)", "route": "oral", "formulary_tier": 3, "prior_auth_required": true, "supply_days_per_fill": 30, "monitoring": "stub-monitor-1 (invented label)", "interaction_flags": ["renavex", "cardiomyst"]}
[5] MODEL_TURN (turn 2) model response, turn 2 [2.31 s · 784 in / 85 out tok · native]
According to the MIS 752 Lab 6 formulary lookup, **zephadril** is on **formulary tier 3** and **does require prior authorisation**. (Source: `get_drug_info` tool result.)
[6] ANSWER (turn 2) the model's final answer [native]
According to the MIS 752 Lab 6 formulary lookup, **zephadril** is on **formulary tier 3** and **does require prior authorisation**. (Source: `get_drug_info` tool result.)
----------------------------------------------------------------------------
PATH PROVENANCE: native
The model emitted a native OpenAI-format tool_calls entry. The API's own
function calling worked, so the prompt protocol was never needed.
TOOLS CALLED: get_drug_info
============================================================================
API calls 5 live + 5 from cache + 0 failed | spend $0.000000
Fallback path: strict JSON in the prompt
The notebook also forced the universal fallback path. The model returned a JSON object requesting get_drug_info, received the same Python result, then returned a JSON final_answer. In this saved run the forced prompt path used 992 prompt tokens versus 1,381 in the auto/native run (−389, or −28%). It was much slower in wall-clock time on that particular free-model run, so the important lesson is to measure rather than assume.
Open Step 5 output (fallback, forecast, sweep, token envelope)
==========================================================================
KNOB 1 · FORCE_PROMPT_PATH = True
==========================================================================
path used : prompt-json
prompt tokens : 992 (the auto-detected run used 1381)
token overhead : -389 tokens (-28%)
tools called : ['get_drug_info']
============================================================================
TRACE · drug_tier-forced
tools ON · path used: prompt-json
model openrouter/free via openrouter · 2 model call(s) (0 from cache) · 131.8 s · 992 in / 235 out tokens · $0.000000
============================================================================
[1] PROMPT (turn 0) the prompt you sent [prompt-json]
system : You are a hospital formulary assistant used in a graduate analytics class (MIS 752, Lab 6). You know NOTHING about the drugs or the arrival forecasts in this conversation except what your tools return. Rules: (1) if a tool can answer, call it instead of guessing; (2) never invent a drug fact, a formulary tier, a drug class or a number; (3) if a tool says there is no record, say plainly that there is no record; (4) answer in at most four sentences and name the tool result you are relying on.
TOOL PROTOCOL. You may call these tools:
- get_drug_info(name: string)
Look up one drug in the MIS ...[+649 chars]
user : What is the formulary tier for zephadril, and does it require prior authorisation?
[2] MODEL_TURN (turn 1) model response, turn 1 [130.65 s · 391 in / 100 out tok · prompt-json]
{"tool": "get_drug_info", "arguments": {"name": "zephadril"}}
[3] TOOL_CALL (turn 1) the model asked for a tool [prompt-json]
name : get_drug_info
arguments : {"name": "zephadril"}
raw : {"tool": "get_drug_info", "arguments": {"name": "zephadril"}}
[4] TOOL_RESULT (turn 1) your Python function answered [0.00 s · prompt-json]
your function returned, in 0.04 ms:
{"drug": "zephadril", "found": true, "synthetic": true, "warning": "SYNTHETIC TEACHING STUB -- these drug names are fictional and every value this stub returns was invented for MIS 752 Lab 6. Inside this lab the stub's answer is authoritative; outside it, nothing here is real. Never use for a clinical decision.", "source": "MIS 752 Lab 6 invented formulary v1", "drug_class": "invented class 7A (teaching placeholder)", "route": "oral", "formulary_tier": 3, "prior_auth_required": true, "supply_days_per_fill": 30, "monitoring": "stub-monitor-1 (invented label)", "interaction_flags": ["renavex", "cardiomyst"]}
[5] MODEL_TURN (turn 2) model response, turn 2 [1.20 s · 601 in / 135 out tok · prompt-json]
{"final_answer": "According to the MIS 752 Lab 6 formulary record for zephadril returned by get_drug_info, this drug is in formulary tier 3 and prior authorization is required (prior_auth_required: true). The record also notes it is an oral medication supplied in 30-day fills. Note that this source is a synthetic teaching stub with fictional values invented for the lab, so these results should not be used for actual clinical decisions."}
[6] ANSWER (turn 2) the model's final answer [prompt-json]
According to the MIS 752 Lab 6 formulary record for zephadril returned by get_drug_info, this drug is in formulary tier 3 and prior authorization is required (prior_auth_required: true). The record also notes it is an oral medication supplied in 30-day fills. Note that this source is a synthetic teaching stub with fictional values invented for the lab, so these results should not be used for actual clinical decisions.
----------------------------------------------------------------------------
PATH PROVENANCE: prompt-json
The prompt-based JSON protocol was used throughout (you forced it, or the
native path was never attempted).
TOOLS CALLED: get_drug_info
============================================================================
==========================================================================
KNOB 2 · the second tool · ed_surge
==========================================================================
the model chose: forecast_ed_arrivals
the question needed: forecast_ed_arrivals
Correct routing. Note what did the work: the DOCSTRING, not the code. The
model never saw forecast_ed_arrivals' body, only its name and description.
============================================================================
TRACE · ed_surge
tools ON · path used: native
model openrouter/free via openrouter · 2 model call(s) (2 from cache) · 0.0 s · 1847 in / 303 out tokens · $0.000000
============================================================================
[1] PROMPT (turn 0) the prompt you sent [native]
system : You are a hospital formulary assistant used in a graduate analytics class (MIS 752, Lab 6). You know NOTHING about the drugs or the arrival forecasts in this conversation except what your tools return. Rules: (1) if a tool can answer, call it instead of guessing; (2) never invent a drug fact, a formulary tier, a drug class or a number; (3) if a tool says there is no record, say plainly that there is no record; (4) answer in at most four sentences and name the tool result you are relying on.
user : Starting 2026-10-05, forecast ED arrivals for the next 7 days. Which day is the peak and how many days are flagged as a surge?
[2] MODEL_TURN (turn 1) model response, turn 1 [1.35 s · 615 in / 111 out tok · native]
[3] TOOL_CALL (turn 1) the model asked for a tool [native]
name : forecast_ed_arrivals
arguments : {"shift_date": "2026-10-05", "horizon_days": 7}
raw : {"shift_date": "2026-10-05", "horizon_days": 7}
[4] TOOL_RESULT (turn 1) your Python function answered [0.00 s · native]
your function returned, in 0.28 ms:
{"synthetic": true, "warning": "SYNTHETIC TEACHING STUB -- this forecast was invented for MIS 752 Lab 6 by a fixed formula, not fitted to any hospital's data. Within this lab the numbers are the authoritative answer; outside it, nothing here is real. Never use for staffing.", "model": "stub seasonal-naive (invented for teaching)", "start": "2026-10-05", "horizon_days": 7, "peak_day": "2026-10-10", "peak_arrivals": 243, "surge_days": 3, "days": [{"date": "2026-10-05", "weekday": "Mon", "predicted_arrivals": 201, "ci_low": 179, "ci_high": 223, "staffing_flag": "surge"}, {"date": "2026-10-06", "weekday": "Tue", "predicted_arrivals": 199, "ci_low": 178, "ci_high": 220, "staffing_flag": "normal"}, {"date": "2026-10-07", "weekday": "Wed", "predicted_arrivals": 181, "ci_low": 162, "ci_high": 200, "staffing_flag": "normal"}, {"date": "2026-10-08", "weekday": "Thu", "predicted_arrivals": 185, "ci ...[+435 chars]
[5] MODEL_TURN (turn 2) model response, turn 2 [2.26 s · 1232 in / 192 out tok · native]
Per the `forecast_ed_arrivals` result, the peak is **Saturday 2026-10-10 with 243 predicted arrivals**. **Three days** are flagged as surges: **2026-10-05, 2026-10-09, and 2026-10-10**. The remaining days (Oct 6, 7, 8, 11) are flagged normal. This is the MIS 752 Lab 6 synthetic teaching forecast, as stated in the tool result.
[6] ANSWER (turn 2) the model's final answer [native]
Per the `forecast_ed_arrivals` result, the peak is **Saturday 2026-10-10 with 243 predicted arrivals**. **Three days** are flagged as surges: **2026-10-05, 2026-10-09, and 2026-10-10**. The remaining days (Oct 6, 7, 8, 11) are flagged normal. This is the MIS 752 Lab 6 synthetic teaching forecast, as stated in the tool result.
----------------------------------------------------------------------------
PATH PROVENANCE: native
The model emitted a native OpenAI-format tool_calls entry. The API's own
function calling worked, so the prompt protocol was never needed.
TOOLS CALLED: forecast_ed_arrivals
============================================================================
==========================================================================
THE SWEEP · 4 questions x 2 conditions, scored as in Step 4
==========================================================================
drug_tier on 0.50 (native) · off 0.00 · delta +0.50
drug_flags on 0.75 (native) · off 0.00 · delta +0.75
drug_absent on 0.43 (native) · off 0.00 · delta +0.43
ed_surge on 0.67 (native) · off 0.00 · delta +0.67
GROUNDING BY QUESTION (share of tool-only facts present in the answer)
condition tools off tools on delta
question tool
drug_absent get_drug_info 0.0 0.43 0.43
drug_flags get_drug_info 0.0 0.75 0.75
drug_tier get_drug_info 0.0 0.50 0.50
ed_surge forecast_ed_arrivals 0.0 0.67 0.67
tools ON scored higher on 4 of 4 comparable questions, lower on 0, tied on 0.
with tools OFF the model made specific formulary claims in 0 of 4 answers, and admitted it had no record in 3.
and in 2 of 4 it NARRATED A TOOL RESULT it never got.
That last number is the one to take to a governance committee. An invented
refusal and an honest one read identically; only the trace separates them.
None of those numbers is a bug. Together they are the risk assessment.
Reading `confident_claims` carefully: a GROUNDED answer usually makes MORE
specific claims than an ungrounded one, because it has real facts to assert.
The column is not a badness score. The question is never 'how many claims?' but
'can each claim be traced to a tool result?' -- and only the trace answers that.
PER-TURN ENVELOPE (every tools-ON model call, in order)
question path turn prompt_tokens completion_tokens seconds
drug_tier native 1 597 69 0.92
drug_tier native 2 784 85 2.31
drug_tier-forced prompt-json 1 391 100 130.65
drug_tier-forced prompt-json 2 601 135 1.20
ed_surge native 1 615 111 1.35
ed_surge native 2 1232 192 2.26
drug_flags native 1 597 89 0.91
drug_flags native 2 783 133 1.22
drug_absent native 1 565 52 1.63
drug_absent native 2 831 144 2.32
prompt tokens ran 391 to 1232 across 10 model calls -- they GROW
with every turn, because the transcript so far is re-sent each time. That is
why an agent that loops 20 times costs far more than 20x one question.
calls this cell: 2 live, 9 from cache
API calls 3 live + 13 from cache + 0 failed | spend $0.000000
Second tool: ED arrival forecast
The model correctly routed the forecasting question to forecast_ed_arrivals. For the synthetic seven-day horizon starting 2026-10-05, the tool returned a peak of 243 arrivals on 2026-10-10 and 3 surge days. The routing decision was driven by the tool name, docstring, and JSON schema—not by access to the Python function body.
Custom extension: hospital bed capacity
This extends the same contract to a student-defined capability: typed arguments, an LLM-readable docstring, a JSON-string result, registration in the tool registry, and prompt-protocol regeneration.
Tool calling vs. MCP
Tool calling
The mechanism: the model asks for a function with structured arguments; application code executes it; the result is returned to the model; the model answers.
MCP framing
The interoperability layer: tools can be advertised with names, descriptions, and input schemas and invoked through a standardized protocol. The lab prints the registry in an MCP-style tools/list envelope but does not run an MCP server.
Open the notebook's MCP comparison output
What you passed the model this whole lab -- openai_tools()
--------------------------------------------------------------------------
{
"type": "function",
"function": {
"name": "get_drug_info",
"description": "Look up one drug in the MIS 752 Lab 6 SYNTHETIC formulary and return its record.",
"parameters": {
"type": "object",
"properties": {
"name": {
"type": "string",
"description": "Drug name to look up, e.g. 'zephadril'."
}
},
"required": [
"name"
]
}
}
}
What an MCP server would advertise for the SAME registry -- tools/list (protocol 2026-07-28)
--------------------------------------------------------------------------
{
"jsonrpc": "2.0",
"id": 1,
"result": {
"tools": [
{
"name": "get_drug_info",
"description": "Look up one drug in the MIS 752 Lab 6 SYNTHETIC formulary and return its record.",
"inputSchema": {
"type": "object",
"properties": {
"name": {
"type": "string",
"description": "Drug name to look up, e.g. 'zephadril'."
}
},
"required": [
"name"
]
}
},
{
"name": "forecast_ed_arrivals",
"description": "Forecast daily emergency-department arrivals from a start date, N days ahead.",
"inputSchema": {
"type": "object",
"properties": {
"shift_date": {
"type": "string",
"description": "First day to forecast, ISO 'YYYY-MM-DD'."
},
"horizon_days": {
"type": "integer",
"minimum": 1,
"maximum": 14,
"description": "How many days ahead to for
And the call your loop already makes, in MCP's envelope -- tools/call
--------------------------------------------------------------------------
{
"jsonrpc": "2.0",
"id": 2,
"method": "tools/call",
"params": {
"name": "get_drug_info",
"arguments": {
"name": "zephadril"
}
}
}
3 tools, described ONCE, rendered into both formats. That is the
whole relationship: the mechanism you built is the part that does not change; MCP
standardises how the description and the result travel.
This lab does NOT run an MCP server. To serve these tools for real, see the `mcp`
Python SDK at https://modelcontextprotocol.io -- your dispatch() is already most of
the way to a tools/call handler.
Reflection
Reproducibility & limitations
This page is a static snapshot of one notebook run. Free-model latency and outputs can vary. The grounding rubric is a teaching metric, not a clinical evaluation. All domain data are synthetic, there is no PHI, and the project should not be used for medical decisions or staffing.
The original notebook is included in this Space package as Lab06_Order_The_Lab_Openrouter_SDahrouch.ipynb for reproducibility.