MIS 752 · Lab 6 · Function / Tool Calling

Order the Lab: teaching a model to reach outside itself

A publishable trace of an LLM using structured tools instead of guessing: native function calling, a strict-JSON fallback, tools-on/tools-off grounding, a custom hospital-capacity tool, and the same registry expressed in MCP-shaped metadata.

OpenRouter / OpenAI-format APINative tool callsPrompt-JSON fallbackMCP concepts$0 measured spend
Synthetic teaching project. Every drug name, formulary value, interaction flag, forecast, and capacity value shown here is invented for MIS 752. Nothing on this page is clinical guidance.

The 30-second read

Native
primary tool-calling path
4 / 4
questions scored higher with tools on
0.00
tools-off grounding across the sweep
$0
measured spend in the saved run
Headline finding: when the model could call the synthetic lookup/forecast tools, tool-only facts appeared in its answers. With tools disabled, those facts disappeared. In two tools-off answers, the model still narrated a “tool result” it never actually received—showing why the trace matters more than confident prose.

Grounding: tools on vs. tools off

Grounding is the share of question-specific facts that could only have come from the synthetic tool data. The saved sweep scored every tools-off answer at 0.00.

drug_tier
0.50
drug_flags
0.75
drug_absent
0.43
ed_surge
0.67

Bars show tools-on grounding. Tools-off grounding was 0.00 for all four questions.

QuestionTools onTools offDelta
drug_tier0.500.00+0.50
drug_flags0.750.00+0.75
drug_absent0.430.00+0.43
ed_surge0.670.00+0.67

Native tool call: formulary lookup

Question

What is the formulary tier for zephadril, and does it require prior authorisation?

Tool selected

get_drug_info(name="zephadril")

Grounded answer

Zephadril is synthetic formulary tier 3 and does require prior authorisation.

Why the trace matters

The model emitted a native structured tool call, Python returned the synthetic JSON record, and the final answer cited that result. The evidence chain is visible rather than inferred from the answer.

Saved run: 2 model calls, 1,381 prompt tokens, 154 completion tokens.

Open the saved native trace
--------------------------------------------------------------------------
COMPARE  ·  run_tool_trace   (same question, same model, same cache)
  agree    path_used
  agree    stage sequence
  agree    tools used
  agree    produced an answer
  A stage-sequence difference is usually one missing make_stage() call; a
  path_used difference usually means you set it from what you INTENDED rather
  than from what happened. Both matter: the trace is the evidence.
--------------------------------------------------------------------------

Steps 4-5 will use the INSTRUCTOR loop. To run them on your own, change the line
above to:  RUN_TRACE = run_tool_trace

============================================================================
 TRACE  ·  drug_tier
 tools ON   ·   path used: native
 model openrouter/free via openrouter   ·   2 model call(s) (2 from cache)   ·   0.0 s   ·   1381 in / 154 out tokens   ·   $0.000000
============================================================================
[1] PROMPT  (turn 0)  the prompt you sent   [native]
      system : You are a hospital formulary assistant used in a graduate analytics class (MIS 752, Lab 6). You know NOTHING about the drugs or the arrival forecasts in this conversation except what your tools return. Rules: (1) if a tool can answer, call it instead of guessing; (2) never invent a drug fact, a formulary tier, a drug class or a number; (3) if a tool says there is no record, say plainly that there is no record; (4) answer in at most four sentences and name the tool result you are relying on.
      user   : What is the formulary tier for zephadril, and does it require prior authorisation?

[2] MODEL_TURN  (turn 1)  model response, turn 1   [0.92 s · 597 in / 69 out tok · native]
      

[3] TOOL_CALL  (turn 1)  the model asked for a tool   [native]
      name      : get_drug_info
      arguments : {"name": "zephadril"}
      raw       : {"name": "zephadril"}

[4] TOOL_RESULT  (turn 1)  your Python function answered   [0.00 s · native]
      your function returned, in 0.03 ms:
      {"drug": "zephadril", "found": true, "synthetic": true, "warning": "SYNTHETIC TEACHING STUB -- these drug names are fictional and every value this stub returns was invented for MIS 752 Lab 6. Inside this lab the stub's answer is authoritative; outside it, nothing here is real. Never use for a clinical decision.", "source": "MIS 752 Lab 6 invented formulary v1", "drug_class": "invented class 7A (teaching placeholder)", "route": "oral", "formulary_tier": 3, "prior_auth_required": true, "supply_days_per_fill": 30, "monitoring": "stub-monitor-1 (invented label)", "interaction_flags": ["renavex", "cardiomyst"]}

[5] MODEL_TURN  (turn 2)  model response, turn 2   [2.31 s · 784 in / 85 out tok · native]
      According to the MIS 752 Lab 6 formulary lookup, **zephadril** is on **formulary tier 3** and **does require prior authorisation**. (Source: `get_drug_info` tool result.)

[6] ANSWER  (turn 2)  the model's final answer   [native]
      According to the MIS 752 Lab 6 formulary lookup, **zephadril** is on **formulary tier 3** and **does require prior authorisation**. (Source: `get_drug_info` tool result.)

----------------------------------------------------------------------------
 PATH PROVENANCE: native
   The model emitted a native OpenAI-format tool_calls entry. The API's own
   function calling worked, so the prompt protocol was never needed.
 TOOLS CALLED: get_drug_info
============================================================================

  API calls 5 live + 5 from cache + 0 failed  |  spend $0.000000

Fallback path: strict JSON in the prompt

The notebook also forced the universal fallback path. The model returned a JSON object requesting get_drug_info, received the same Python result, then returned a JSON final_answer. In this saved run the forced prompt path used 992 prompt tokens versus 1,381 in the auto/native run (−389, or −28%). It was much slower in wall-clock time on that particular free-model run, so the important lesson is to measure rather than assume.

Open Step 5 output (fallback, forecast, sweep, token envelope)
==========================================================================
 KNOB 1 · FORCE_PROMPT_PATH = True
==========================================================================
  path used      : prompt-json
  prompt tokens  : 992 (the auto-detected run used 1381)
  token overhead : -389 tokens (-28%)
  tools called   : ['get_drug_info']

============================================================================
 TRACE  ·  drug_tier-forced
 tools ON   ·   path used: prompt-json
 model openrouter/free via openrouter   ·   2 model call(s) (0 from cache)   ·   131.8 s   ·   992 in / 235 out tokens   ·   $0.000000
============================================================================
[1] PROMPT  (turn 0)  the prompt you sent   [prompt-json]
      system : You are a hospital formulary assistant used in a graduate analytics class (MIS 752, Lab 6). You know NOTHING about the drugs or the arrival forecasts in this conversation except what your tools return. Rules: (1) if a tool can answer, call it instead of guessing; (2) never invent a drug fact, a formulary tier, a drug class or a number; (3) if a tool says there is no record, say plainly that there is no record; (4) answer in at most four sentences and name the tool result you are relying on.

TOOL PROTOCOL. You may call these tools:
- get_drug_info(name: string)
    Look up one drug in the MIS  ...[+649 chars]
      user   : What is the formulary tier for zephadril, and does it require prior authorisation?

[2] MODEL_TURN  (turn 1)  model response, turn 1   [130.65 s · 391 in / 100 out tok · prompt-json]
      {"tool": "get_drug_info", "arguments": {"name": "zephadril"}}

[3] TOOL_CALL  (turn 1)  the model asked for a tool   [prompt-json]
      name      : get_drug_info
      arguments : {"name": "zephadril"}
      raw       : {"tool": "get_drug_info", "arguments": {"name": "zephadril"}}

[4] TOOL_RESULT  (turn 1)  your Python function answered   [0.00 s · prompt-json]
      your function returned, in 0.04 ms:
      {"drug": "zephadril", "found": true, "synthetic": true, "warning": "SYNTHETIC TEACHING STUB -- these drug names are fictional and every value this stub returns was invented for MIS 752 Lab 6. Inside this lab the stub's answer is authoritative; outside it, nothing here is real. Never use for a clinical decision.", "source": "MIS 752 Lab 6 invented formulary v1", "drug_class": "invented class 7A (teaching placeholder)", "route": "oral", "formulary_tier": 3, "prior_auth_required": true, "supply_days_per_fill": 30, "monitoring": "stub-monitor-1 (invented label)", "interaction_flags": ["renavex", "cardiomyst"]}

[5] MODEL_TURN  (turn 2)  model response, turn 2   [1.20 s · 601 in / 135 out tok · prompt-json]
      {"final_answer": "According to the MIS 752 Lab 6 formulary record for zephadril returned by get_drug_info, this drug is in formulary tier 3 and prior authorization is required (prior_auth_required: true). The record also notes it is an oral medication supplied in 30-day fills. Note that this source is a synthetic teaching stub with fictional values invented for the lab, so these results should not be used for actual clinical decisions."}

[6] ANSWER  (turn 2)  the model's final answer   [prompt-json]
      According to the MIS 752 Lab 6 formulary record for zephadril returned by get_drug_info, this drug is in formulary tier 3 and prior authorization is required (prior_auth_required: true). The record also notes it is an oral medication supplied in 30-day fills. Note that this source is a synthetic teaching stub with fictional values invented for the lab, so these results should not be used for actual clinical decisions.

----------------------------------------------------------------------------
 PATH PROVENANCE: prompt-json
   The prompt-based JSON protocol was used throughout (you forced it, or the
   native path was never attempted).
 TOOLS CALLED: get_drug_info
============================================================================

==========================================================================
 KNOB 2 · the second tool  ·  ed_surge
==========================================================================
  the model chose: forecast_ed_arrivals
  the question needed: forecast_ed_arrivals
  Correct routing. Note what did the work: the DOCSTRING, not the code. The
  model never saw forecast_ed_arrivals' body, only its name and description.

============================================================================
 TRACE  ·  ed_surge
 tools ON   ·   path used: native
 model openrouter/free via openrouter   ·   2 model call(s) (2 from cache)   ·   0.0 s   ·   1847 in / 303 out tokens   ·   $0.000000
============================================================================
[1] PROMPT  (turn 0)  the prompt you sent   [native]
      system : You are a hospital formulary assistant used in a graduate analytics class (MIS 752, Lab 6). You know NOTHING about the drugs or the arrival forecasts in this conversation except what your tools return. Rules: (1) if a tool can answer, call it instead of guessing; (2) never invent a drug fact, a formulary tier, a drug class or a number; (3) if a tool says there is no record, say plainly that there is no record; (4) answer in at most four sentences and name the tool result you are relying on.
      user   : Starting 2026-10-05, forecast ED arrivals for the next 7 days. Which day is the peak and how many days are flagged as a surge?

[2] MODEL_TURN  (turn 1)  model response, turn 1   [1.35 s · 615 in / 111 out tok · native]
      

[3] TOOL_CALL  (turn 1)  the model asked for a tool   [native]
      name      : forecast_ed_arrivals
      arguments : {"shift_date": "2026-10-05", "horizon_days": 7}
      raw       : {"shift_date": "2026-10-05", "horizon_days": 7}

[4] TOOL_RESULT  (turn 1)  your Python function answered   [0.00 s · native]
      your function returned, in 0.28 ms:
      {"synthetic": true, "warning": "SYNTHETIC TEACHING STUB -- this forecast was invented for MIS 752 Lab 6 by a fixed formula, not fitted to any hospital's data. Within this lab the numbers are the authoritative answer; outside it, nothing here is real. Never use for staffing.", "model": "stub seasonal-naive (invented for teaching)", "start": "2026-10-05", "horizon_days": 7, "peak_day": "2026-10-10", "peak_arrivals": 243, "surge_days": 3, "days": [{"date": "2026-10-05", "weekday": "Mon", "predicted_arrivals": 201, "ci_low": 179, "ci_high": 223, "staffing_flag": "surge"}, {"date": "2026-10-06", "weekday": "Tue", "predicted_arrivals": 199, "ci_low": 178, "ci_high": 220, "staffing_flag": "normal"}, {"date": "2026-10-07", "weekday": "Wed", "predicted_arrivals": 181, "ci_low": 162, "ci_high": 200, "staffing_flag": "normal"}, {"date": "2026-10-08", "weekday": "Thu", "predicted_arrivals": 185, "ci ...[+435 chars]

[5] MODEL_TURN  (turn 2)  model response, turn 2   [2.26 s · 1232 in / 192 out tok · native]
      Per the `forecast_ed_arrivals` result, the peak is **Saturday 2026-10-10 with 243 predicted arrivals**. **Three days** are flagged as surges: **2026-10-05, 2026-10-09, and 2026-10-10**. The remaining days (Oct 6, 7, 8, 11) are flagged normal. This is the MIS 752 Lab 6 synthetic teaching forecast, as stated in the tool result.

[6] ANSWER  (turn 2)  the model's final answer   [native]
      Per the `forecast_ed_arrivals` result, the peak is **Saturday 2026-10-10 with 243 predicted arrivals**. **Three days** are flagged as surges: **2026-10-05, 2026-10-09, and 2026-10-10**. The remaining days (Oct 6, 7, 8, 11) are flagged normal. This is the MIS 752 Lab 6 synthetic teaching forecast, as stated in the tool result.

----------------------------------------------------------------------------
 PATH PROVENANCE: native
   The model emitted a native OpenAI-format tool_calls entry. The API's own
   function calling worked, so the prompt protocol was never needed.
 TOOLS CALLED: forecast_ed_arrivals
============================================================================

==========================================================================
 THE SWEEP · 4 questions x 2 conditions, scored as in Step 4
==========================================================================
  drug_tier    on 0.50 (native)   ·   off 0.00   ·   delta +0.50
  drug_flags   on 0.75 (native)   ·   off 0.00   ·   delta +0.75
  drug_absent  on 0.43 (native)   ·   off 0.00   ·   delta +0.43
  ed_surge     on 0.67 (native)   ·   off 0.00   ·   delta +0.67

GROUNDING BY QUESTION (share of tool-only facts present in the answer)
condition                         tools off  tools on  delta
question    tool                                            
drug_absent get_drug_info               0.0      0.43   0.43
drug_flags  get_drug_info               0.0      0.75   0.75
drug_tier   get_drug_info               0.0      0.50   0.50
ed_surge    forecast_ed_arrivals        0.0      0.67   0.67

  tools ON scored higher on 4 of 4 comparable questions, lower on 0, tied on 0.
  with tools OFF the model made specific formulary claims in 0 of 4 answers, and admitted it had no record in 3.
  and in 2 of 4 it NARRATED A TOOL RESULT it never got.
  That last number is the one to take to a governance committee. An invented
  refusal and an honest one read identically; only the trace separates them.
  None of those numbers is a bug. Together they are the risk assessment.

  Reading `confident_claims` carefully: a GROUNDED answer usually makes MORE
  specific claims than an ungrounded one, because it has real facts to assert.
  The column is not a badness score. The question is never 'how many claims?' but
  'can each claim be traced to a tool result?' -- and only the trace answers that.

PER-TURN ENVELOPE (every tools-ON model call, in order)
        question        path  turn  prompt_tokens  completion_tokens  seconds
       drug_tier      native     1            597                 69     0.92
       drug_tier      native     2            784                 85     2.31
drug_tier-forced prompt-json     1            391                100   130.65
drug_tier-forced prompt-json     2            601                135     1.20
        ed_surge      native     1            615                111     1.35
        ed_surge      native     2           1232                192     2.26
      drug_flags      native     1            597                 89     0.91
      drug_flags      native     2            783                133     1.22
     drug_absent      native     1            565                 52     1.63
     drug_absent      native     2            831                144     2.32

  prompt tokens ran 391 to 1232 across 10 model calls -- they GROW
  with every turn, because the transcript so far is re-sent each time. That is
  why an agent that loops 20 times costs far more than 20x one question.

  calls this cell: 2 live, 9 from cache
  API calls 3 live + 13 from cache + 0 failed  |  spend $0.000000

Second tool: ED arrival forecast

The model correctly routed the forecasting question to forecast_ed_arrivals. For the synthetic seven-day horizon starting 2026-10-05, the tool returned a peak of 243 arrivals on 2026-10-10 and 3 surge days. The routing decision was driven by the tool name, docstring, and JSON schema—not by access to the Python function body.

Custom extension: hospital bed capacity

For my experiment, I created a get_unit_capacity tool that returns synthetic hospital bed-capacity information and registered it with the same tool system used throughout the lab. I rebuilt PROMPT_PROTOCOL so the prompt-based path could recognize the new tool, then tested it with questions containing facts that could only come from my tool. Comparing the tools-on and tools-off results showed how tool use affects the grounding of the model's answers. I used OpenAI's ChatGPT to help me understand the assignment, develop and troubleshoot the custom tool code, and write up my findings.

This extends the same contract to a student-defined capability: typed arguments, an LLM-readable docstring, a JSON-string result, registration in the tool registry, and prompt-protocol regeneration.

Tool calling vs. MCP

Tool calling

The mechanism: the model asks for a function with structured arguments; application code executes it; the result is returned to the model; the model answers.

MCP framing

The interoperability layer: tools can be advertised with names, descriptions, and input schemas and invoked through a standardized protocol. The lab prints the registry in an MCP-style tools/list envelope but does not run an MCP server.

Open the notebook's MCP comparison output
What you passed the model this whole lab  --  openai_tools()
--------------------------------------------------------------------------
{
 "type": "function",
 "function": {
  "name": "get_drug_info",
  "description": "Look up one drug in the MIS 752 Lab 6 SYNTHETIC formulary and return its record.",
  "parameters": {
   "type": "object",
   "properties": {
    "name": {
     "type": "string",
     "description": "Drug name to look up, e.g. 'zephadril'."
    }
   },
   "required": [
    "name"
   ]
  }
 }
}

What an MCP server would advertise for the SAME registry  --  tools/list (protocol 2026-07-28)
--------------------------------------------------------------------------
{
 "jsonrpc": "2.0",
 "id": 1,
 "result": {
  "tools": [
   {
    "name": "get_drug_info",
    "description": "Look up one drug in the MIS 752 Lab 6 SYNTHETIC formulary and return its record.",
    "inputSchema": {
     "type": "object",
     "properties": {
      "name": {
       "type": "string",
       "description": "Drug name to look up, e.g. 'zephadril'."
      }
     },
     "required": [
      "name"
     ]
    }
   },
   {
    "name": "forecast_ed_arrivals",
    "description": "Forecast daily emergency-department arrivals from a start date, N days ahead.",
    "inputSchema": {
     "type": "object",
     "properties": {
      "shift_date": {
       "type": "string",
       "description": "First day to forecast, ISO 'YYYY-MM-DD'."
      },
      "horizon_days": {
       "type": "integer",
       "minimum": 1,
       "maximum": 14,
       "description": "How many days ahead to for

And the call your loop already makes, in MCP's envelope  --  tools/call
--------------------------------------------------------------------------
{
 "jsonrpc": "2.0",
 "id": 2,
 "method": "tools/call",
 "params": {
  "name": "get_drug_info",
  "arguments": {
   "name": "zephadril"
  }
 }
}

3 tools, described ONCE, rendered into both formats. That is the
whole relationship: the mechanism you built is the part that does not change; MCP
standardises how the description and the result travel.

This lab does NOT run an MCP server. To serve these tools for real, see the `mcp`
Python SDK at https://modelcontextprotocol.io -- your dispatch() is already most of
the way to a tools/call handler.

Reflection

My model primarily used the native tool-calling path, with the trace reporting `path_used: native`. In the successful runs, the model correctly called `get_drug_info` and `forecast_ed_arrivals`, so I did not observe the native path failing and switching to the prompt-based fallback. With tools turned off, some answers had no usable result, and two runs were truncated; this is important in a clinical setting because a model without access to the actual record could guess, omit information, or present an unsupported claim as fact. I did not successfully measure the cost of the prompt-based fallback because `FORCE_PROMPT_PATH` was False; my comparison therefore showed 1,381 prompt tokens for both runs and +0 measured overhead. Outside healthcare, function calling would be useful in banking, where a model should retrieve an account balance or transaction from an authorized system instead of answering from memory. I used OpenAI’s ChatGPT for help understanding and troubleshooting this lab.

Reproducibility & limitations

This page is a static snapshot of one notebook run. Free-model latency and outputs can vary. The grounding rubric is a teaching metric, not a clinical evaluation. All domain data are synthetic, there is no PHI, and the project should not be used for medical decisions or staffing.

The original notebook is included in this Space package as Lab06_Order_The_Lab_Openrouter_SDahrouch.ipynb for reproducibility.