🧭 Tool Schema Calibration#
gpt-4o-mini · temperature 0 · tool_choice="required" · names × descriptions 2x2. Open a topic to see the idea, the request path and the function calls behind the demo, then read the complete Python source file by file.
How It Works#
The idea behind the demo, the request it sends and the function calls that answer it.
Concept#
The model is offered three weather tools — current, forecast, and history — that take the same parameters. The only things that differ are each tool's name and description. Crossing two naming styles with two description styles gives four conditions, and the demo measures how often the model picks the right tool in each.
Your own question is routed under all four conditions so you can see the choices side by side, then 12 labelled queries (4 per tool) fill in the accuracy matrix.
Theory & Concepts#
1. The 2x2 design
- A · descriptive names + loose descriptions —
get_weather_forecast/ "Get weather data for a city." - B · descriptive names + tight descriptions — names plus "Use only for 'tomorrow', 'next week'…"
- C · opaque names + loose descriptions —
weather_service_b/ vague prose. No real signal at all. - D · opaque names + tight descriptions — the description is the only signal.
2. Why the names had to vary too
An earlier version varied only the descriptions while the names stayed
get_weather_current/_forecast/_history. Those names already encode the distinction,
so the model routed on the name and scored near 100% either way. Crossing names with descriptions
separates the two signals.
3. Forced tool choice
tool_choice="required" makes the model pick one of the three tools. That isolates
"which tool" from "whether to call a tool at all" — two different bugs with different fixes.
Temperature 0 keeps the schema text as the only variable.
4. The practical lesson
Descriptions look optional right up until someone renames a tool, or you mount an MCP server whose
tools are called query, search, and execute. Condition C
is what that feels like.
Request flow#
Code flow#
question] -->|POST /routing| B[app.py
routing_route] B -->|question| C[tool_routing.py
run_routing] C -->|52 jobs| D[config.py
parallel_map] D -->|query, name style, desc style| E[select_tool] E -->|styles| G[build_condition] G -->|tools + lookup| E E -->|tools, tool_choice=required| F[OpenAI API
gpt-4o-mini] F -->|function name| E E -->|canonical slot| C C -->|text report| B B -->|JSON result| A
Source Code#
Every Python file this demo runs, complete and unedited: the feature code first, then the shared Flask routes.
tool_routing.py#
Tool Schema Calibration: which part of a tool schema does the model route on?
"""Tool Schema Calibration: which part of a tool schema does the model route on?
Adapted from ``study/09-ai-reliability/tool-schema-calibration.py``. Three weather
tools are offered under a 2x2 of conditions (descriptive vs opaque names x
loose vs tight descriptions). ``tool_choice="required"`` forces a pick so
"which tool" is isolated from "whether to call a tool at all".
The web demo routes the visitor's own question under all four conditions, then
scores a fixed labelled sample to fill in the accuracy matrix.
"""
from config import CHAT_MODEL, bar, get_openai_client, parallel_map
TEMPERATURE = 0
MAX_QUERY_CHARS = 300
CANONICAL = ("current", "forecast", "history")
NAME_SETS = {
"descriptive": {
"current": "get_weather_current",
"forecast": "get_weather_forecast",
"history": "get_weather_history",
},
# plausible but uninformative - the shape real MCP servers often ship with
"opaque": {
"current": "weather_service_a",
"forecast": "weather_service_b",
"history": "weather_service_c",
},
}
DESCRIPTION_SETS = {
"loose": {
"current": "Get weather information for a location.",
"forecast": "Get weather data for a city.",
"history": "Fetch weather records for a specific area.",
},
"tight": {
"current": (
"Get CURRENT, REAL-TIME weather conditions for a specific location. "
"Use only for queries about 'now', 'today', or current status."
),
"forecast": (
"Get FUTURE weather predictions and forecasts. Use only for queries "
"about 'tomorrow', 'next week', 'upcoming', or future dates."
),
"history": (
"Fetch HISTORICAL weather records from the past. Use only for queries "
"about yesterday, last year, or specific past dates."
),
},
}
CONDITIONS = [
("A", "descriptive", "loose"),
("B", "descriptive", "tight"),
("C", "opaque", "loose"),
("D", "opaque", "tight"),
]
NAME_CHOICES = ("both", *NAME_SETS)
DESCRIPTION_CHOICES = ("both", *DESCRIPTION_SETS)
# A balanced subset of the study prototype's 50 labelled queries.
SAMPLE_QUERIES = [
("Is it raining in London at the moment?", "current"),
("Show me today's weather for NYC.", "current"),
("Check weather for Rome.", "current"),
("Weather update for Cape Town.", "current"),
("Will it rain next Tuesday in Paris?", "forecast"),
("Give me the 5-day forecast for Sydney.", "forecast"),
("Upcoming weather for Chicago.", "forecast"),
("Weather outlook for Seoul next month.", "forecast"),
("What was the weather like in London yesterday?", "history"),
("Weather records for Tokyo in 1990.", "history"),
("Last week's weather in Toronto.", "history"),
("Was it raining in Seoul three days ago?", "history"),
]
def build_condition(name_style: str, desc_style: str) -> tuple[list, dict]:
"""Return (tools, lookup) where lookup maps emitted names to canonical slots."""
# ① choose the name and description set for this calibration condition
names = NAME_SETS[name_style]
descs = DESCRIPTION_SETS[desc_style]
# ② build OpenAI tool schemas for the same three weather capabilities
tools = [
{
"type": "function",
"function": {
"name": names[c],
"description": descs[c],
"parameters": {
"type": "object",
"properties": {"location": {"type": "string", "description": "The city name."}},
"required": ["location"],
},
},
}
for c in CANONICAL
]
# ③ return the schemas plus a lookup back to the canonical answer labels
return tools, {names[c]: c for c in CANONICAL}
def select_tool(query: str, name_style: str, desc_style: str) -> str:
"""Return the canonical slot the model routed to, '(no call)', or '(error)'."""
# ① build the exact tool menu for this name-description condition
tools, lookup = build_condition(name_style, desc_style)
try:
# ② force the model to choose one tool so routing can be measured directly
response = get_openai_client().chat.completions.create(
model=CHAT_MODEL,
messages=[{"role": "user", "content": query}],
tools=tools,
tool_choice="required",
temperature=TEMPERATURE,
)
# ③ read the selected tool call and map it back to current/forecast/history
calls = response.choices[0].message.tool_calls
if not calls:
return "(no call)"
return lookup.get(calls[0].function.name, calls[0].function.name)
except Exception:
# ④ keep API failures visible as a routing outcome instead of crashing
return "(error)"
def run_routing(query: str, names: str = "both", descriptions: str = "both") -> str:
"""Route ``query`` under the selected conditions and score the labelled sample."""
# ① trim the learner's question and choose the requested 2x2 conditions
query = query.strip()[:MAX_QUERY_CHARS]
conditions = [
c for c in CONDITIONS
if names in ("both", c[1]) and descriptions in ("both", c[2])
]
# ② queue the learner query first, then the labelled sample for each condition
jobs = [(query, n, d) for _, n, d in conditions]
jobs += [(q, n, d) for _, n, d in conditions for q, _ in SAMPLE_QUERIES]
# ③ run all routing decisions in parallel to keep the demo responsive
picks = parallel_map(lambda job: select_tool(*job), jobs)
# ④ split personal picks from sample picks and score each condition
yours, sample = picks[: len(conditions)], picks[len(conditions):]
total = len(SAMPLE_QUERIES)
scores = {}
for i, (label, _, _) in enumerate(conditions):
chunk = sample[i * total:(i + 1) * total]
scores[label] = sum(p == exp for p, (_, exp) in zip(chunk, SAMPLE_QUERIES))
# ⑤ report how the learner's question routed under every selected condition
lines = [f"Your question · {CHAT_MODEL} · temperature {TEMPERATURE} · tool_choice=required"]
for (label, n, d), pick in zip(conditions, yours):
tool_name = NAME_SETS[n].get(pick, pick)
lines.append(f" {label} {n} names + {d} descs → {tool_name} ({pick})")
# ⑥ add the labelled-sample accuracy matrix and lift comparisons
lines += ["", f"Routing accuracy on {total} labelled queries"]
for label, n, d in conditions:
lines.append(f" {label} {n:<11} + {d:<5} {bar(scores[label], total)}")
lifts = [
("Description lift with self-explanatory names", "B", "A"),
("Description lift with opaque names", "D", "C"),
("Name lift with vague descriptions", "A", "C"),
("Name lift with tight descriptions", "B", "D"),
]
lift_lines = [
f"{text}: {scores[hi] - scores[lo]:+d}" for text, hi, lo in lifts if hi in scores and lo in scores
]
if lift_lines:
lines += [""] + lift_lines
# ⑦ explain either how to see the full grid or what the grid shows
lines.append("")
if len(scores) < 4:
lines.append("Tip: set both dropdowns to 'Both' for the full 2x2 and a verdict.")
return "\n".join(lines)
desc_lift_desc = scores["B"] - scores["A"]
desc_lift_opaque = scores["D"] - scores["C"]
if desc_lift_opaque > desc_lift_desc:
lines.append(
"Verdict: descriptions are load-bearing - but only when the names are not "
"already doing the work. Rename tools generically and the description is the "
"only signal left."
)
elif desc_lift_desc > 0 or desc_lift_opaque > 0:
lines.append("Verdict: tightening descriptions improves routing regardless of naming.")
else:
lines.append(
"Verdict: no measurable effect from descriptions on this sample - the model "
"is saturating the task."
)
return "\n".join(lines)
if __name__ == "__main__":
# ① run one forecast-style example when this module is executed directly
print(run_routing("Will it snow in Oslo this weekend?"))
config.py#
Shared configuration: load .env and build the OpenAI client.
"""Shared configuration: load .env and build the OpenAI client.
This module is the single place that knows about secrets and model names.
Every feature module imports from here instead of reading ``os.environ`` or
constructing API clients itself.
"""
import os
from concurrent.futures import ThreadPoolExecutor
from dotenv import load_dotenv
from openai import OpenAI
load_dotenv()
def get_env(name: str, default: str = "") -> str:
"""Return an environment variable, falling back to ``default``."""
return os.environ.get(name, default)
# A NON-reasoning model: reasoning models think internally even when told not
# to, which erases the gaps these demos measure.
CHAT_MODEL = get_env("OPENAI_MODEL", "gpt-4o-mini")
# Upper bound on parallel API calls per request; keeps each demo well under
# the 60 s Nginx proxy timeout without hammering the provider.
MAX_WORKERS = 8
# Temperatures selectable from the UI, keyed by the string the browser sends.
TEMPERATURE_CHOICES = {"0": 0.0, "0.7": 0.7, "1.2": 1.2}
_client = None
def get_openai_client() -> OpenAI:
"""Return a shared OpenAI client built from OPENAI_API_KEY."""
global _client
if _client is None:
# ① read the API key lazily so tests/imports do not require credentials
api_key = get_env("OPENAI_API_KEY")
if not api_key:
raise RuntimeError("OPENAI_API_KEY is not set. Add it to your .env file.")
# ② create one reusable client with bounded timeout and retries
_client = OpenAI(api_key=api_key, timeout=30, max_retries=2)
return _client
def parallel_map(fn, items):
"""Run ``fn`` over ``items`` concurrently, preserving input order."""
with ThreadPoolExecutor(max_workers=MAX_WORKERS) as pool:
return list(pool.map(fn, items))
def bar(correct: int, total: int, width: int = 10) -> str:
"""Render a text accuracy bar with a count and percentage."""
# ① avoid dividing by zero when a filtered comparison has no examples
if total == 0:
return "n/a"
# ② convert the score into filled and empty bar characters
filled = round(correct / total * width)
return f"{'█' * filled}{'░' * (width - filled)} {correct}/{total} ({round(correct / total * 100)}%)"
app.py#
Flask server for AI Reliability Lab: four measured LLM reliability demos.
"""Flask server for AI Reliability Lab: four measured LLM reliability demos.
Architecture notes
------------------
- All routes are attached to a Blueprint (``bp``) instead of directly to
``app``. This lets us register the entire Blueprint under a runtime URL
prefix (``PATH_PREFIX``) without touching individual route strings.
- In local development PATH_PREFIX is empty, so routes are at "/",
"/variance", etc. In production Nginx forwards ``/ai-reliability/...``
traffic to the container and PATH_PREFIX is set to "/ai-reliability".
- flask-cors adds ``Access-Control-Allow-Origin: *`` headers so the HTML
page can call the API even if it is served from a different origin during
development.
"""
import os
from pathlib import Path
from flask import Blueprint, Flask, jsonify, request
from flask_cors import CORS
from config import TEMPERATURE_CHOICES
from cot import PROBLEMS, run_cot
from cot import STRATEGY_CHOICES as COT_STRATEGIES
from rate_limiter import check_rate_limit
from tool_errors import FAIL_ON_CALL_LABELS, PROMPT_CHOICES, run_error_injection
from tool_routing import DESCRIPTION_CHOICES, NAME_CHOICES, run_routing
from variance import STRATEGY_CHOICES as VARIANCE_STRATEGIES
from variance import run_variance
# ---------------------------------------------------------------------------
# Configuration
# ---------------------------------------------------------------------------
# PATH_PREFIX is set by the deployment environment ("/ai-reliability") so
# the app works correctly behind an Nginx location block. Locally it is an
# empty string, which mounts all routes at the root.
PATH_PREFIX = os.environ.get("PATH_PREFIX", "")
# app.py lives in src/python, while index.html, css/, and js/ live in src/.
STATIC_DIR = Path(__file__).resolve().parents[1]
app = Flask(__name__, static_folder=str(STATIC_DIR))
# Allow cross-origin requests from any origin. In production you would
# restrict this to the specific front-end domain.
CORS(app)
# A Blueprint groups related routes. We register it once at the bottom with
# the runtime PATH_PREFIX, avoiding any hardcoded path strings in the routes.
bp = Blueprint("main", __name__)
@bp.before_request
def enforce_rate_limit():
"""Enforce strict 10 requests per hour limit on all POST endpoints."""
# ① apply the limit only to API actions, not static page loads
if request.method == "POST":
# ② ask the shared limiter whether this request should be blocked
blocked, msg, retry_after = check_rate_limit(
request, max_requests=10, window_seconds=3600
)
if blocked:
# ③ return a 429 with retry guidance when the hourly quota is exhausted
resp = jsonify({"error": msg})
resp.status_code = 429
resp.headers["Retry-After"] = str(retry_after)
return resp
# ---------------------------------------------------------------------------
# Routes
# ---------------------------------------------------------------------------
@bp.route("/")
def index():
"""Serve index.html, injecting the correct API base URL for the environment."""
# ① read the static HTML shell from the configured Flask static folder
with open(os.path.join(app.static_folder, "index.html"), encoding="utf-8") as f:
html = f.read()
# ② inject the runtime path prefix so browser fetches target the right API base
# The HTML file ships with 'data-api-base=""' (empty = relative URL, works
# locally). For production we replace it with the actual path prefix so
# all fetch() calls in the browser target the right endpoint.
html = html.replace('data-api-base=""', f'data-api-base="{PATH_PREFIX}"')
# ③ return the modified HTML with an explicit text/html response type
return app.response_class(html, mimetype="text/html")
@bp.route("/css/<path:filename>")
def css(filename):
"""Serve stylesheets from the src/css directory."""
return app.send_static_file(os.path.join("css", filename))
@bp.route("/js/<path:filename>")
def js(filename):
"""Serve scripts from the src/js directory."""
return app.send_static_file(os.path.join("js", filename))
@bp.route("/info/<path:filename>")
def info(filename):
"""Serve the "how this demo works" explainer pages from src/info."""
return app.send_static_file(os.path.join("info", filename))
def read_message() -> str:
"""Return the trimmed ``message`` field from the JSON body, or ''."""
# ① parse JSON leniently so missing or malformed bodies become empty data
data = request.get_json(force=True, silent=True) or {}
# ② normalize the message field into a stripped string for route validation
return str(data.get("message") or "").strip()
def read_choice(name: str, allowed, default: str) -> str | None:
"""Return a dropdown value from the JSON body, ``default`` if absent, or None if not allowed."""
# ① parse JSON leniently and fall back to the route's default choice
data = request.get_json(force=True, silent=True) or {}
value = str(data.get(name) or default)
# ② accept only known UI choices so feature modules receive valid selectors
return value if value in allowed else None
def invalid_choice(name: str, allowed):
return jsonify({"error": f"Invalid {name}. Choose one of: {', '.join(map(str, allowed))}."}), 400
@bp.route("/variance", methods=["POST"])
def variance_route():
"""Run Unconstrained vs Prompt-JSON vs Schema-enforced extraction on a review."""
# ① validate that the learner supplied review text to analyze
message = read_message()
if not message:
return jsonify({"error": "A customer review is required."}), 400
# ② read and validate the selected extraction strategy
strategy = read_choice("strategy", VARIANCE_STRATEGIES, "all")
if strategy is None:
return invalid_choice("strategy", VARIANCE_STRATEGIES)
# ③ read and validate the selected temperature
temperature = read_choice("temperature", TEMPERATURE_CHOICES, "0.7")
if temperature is None:
return invalid_choice("temperature", TEMPERATURE_CHOICES)
try:
# ④ run the variance feature and return its text report as JSON
return jsonify({"result": run_variance(message, strategy, TEMPERATURE_CHOICES[temperature])})
except Exception:
# ⑤ log server-side detail while returning a safe client-facing error
app.logger.exception("variance failed")
return jsonify({"error": "Variance test failed. Please try again later."}), 500
@bp.route("/cot", methods=["POST"])
def cot_route():
"""Run Direct vs Chain-of-Thought prompting on one preset problem."""
# ① read the requested problem key and normalize it for lookup
message = read_message().lower()
if not message:
return jsonify({"error": "A problem name is required."}), 400
if message not in PROBLEMS:
return jsonify({"error": f"Unknown problem. Choose one of: {', '.join(PROBLEMS)}."}), 400
# ② read and validate the selected prompt strategy
strategy = read_choice("strategy", COT_STRATEGIES, "both")
if strategy is None:
return invalid_choice("strategy", COT_STRATEGIES)
# ③ read and validate the selected temperature
temperature = read_choice("temperature", TEMPERATURE_CHOICES, "0.7")
if temperature is None:
return invalid_choice("temperature", TEMPERATURE_CHOICES)
try:
# ④ run the CoT feature and return its text report as JSON
return jsonify({"result": run_cot(message, strategy, TEMPERATURE_CHOICES[temperature])})
except Exception:
# ⑤ log server-side detail while returning a safe client-facing error
app.logger.exception("cot failed")
return jsonify({"error": "CoT comparison failed. Please try again later."}), 500
@bp.route("/routing", methods=["POST"])
def routing_route():
"""Route a weather question under the names x descriptions 2x2."""
# ① validate that the learner supplied a weather-routing question
message = read_message()
if not message:
return jsonify({"error": "A weather question is required."}), 400
# ② read and validate the selected tool-name condition
names = read_choice("names", NAME_CHOICES, "both")
if names is None:
return invalid_choice("names", NAME_CHOICES)
# ③ read and validate the selected tool-description condition
descriptions = read_choice("descriptions", DESCRIPTION_CHOICES, "both")
if descriptions is None:
return invalid_choice("descriptions", DESCRIPTION_CHOICES)
try:
# ④ run the routing feature and return its text report as JSON
return jsonify({"result": run_routing(message, names, descriptions)})
except Exception:
# ⑤ log server-side detail while returning a safe client-facing error
app.logger.exception("routing failed")
return jsonify({"error": "Routing test failed. Please try again later."}), 500
@bp.route("/errors", methods=["POST"])
def errors_route():
"""Run the tool-error injection scenarios on a weather question."""
# ① validate that the learner supplied a weather question for the agent
message = read_message()
if not message:
return jsonify({"error": "A weather question is required."}), 400
# ② read and validate which system-prompt scenario to run
prompt = read_choice("prompt", PROMPT_CHOICES, "both")
if prompt is None:
return invalid_choice("prompt", PROMPT_CHOICES)
# ③ derive valid failure-injection choices from the shared labels
fail_choices = [str(k) for k in FAIL_ON_CALL_LABELS]
fail_on_call = read_choice("fail_on_call", fail_choices, "2")
if fail_on_call is None:
return invalid_choice("fail_on_call", fail_choices)
try:
# ④ run the error-injection feature and return its text report as JSON
return jsonify({"result": run_error_injection(message, prompt, int(fail_on_call))})
except Exception:
# ⑤ log server-side detail while returning a safe client-facing error
app.logger.exception("error injection failed")
return jsonify({"error": "Error-injection test failed. Please try again later."}), 500
# ---------------------------------------------------------------------------
# Blueprint registration + server entry point
# ---------------------------------------------------------------------------
# Register all Blueprint routes under the optional path prefix. This single
# line is the only place where PATH_PREFIX is applied — every route above is
# written as a relative path (e.g. "/variance") and the prefix is prepended here.
app.register_blueprint(bp, url_prefix=PATH_PREFIX)
if __name__ == "__main__":
# Run the development server. 0.0.0.0 makes the app reachable from outside
# the container; port 5000 is mapped to the host port in docker-compose.yml.
app.run(host="0.0.0.0", port=5000)