🎲 Variance & Determinism#
gpt-4o-mini · temperature 0.7 · 3 strategies × 5 runs · strict json.loads() grading. Open a topic to see the idea, the request path and the function calls behind the demo, then read the complete Python source file by file.
How It Works#
The idea behind the demo, the request it sends and the function calls that answer it.
Concept#
The same extraction task — "give me the sentiment and entities of this review" — is sent
15 times: 5 runs each under three strategies. Every response is then handed to a
strict downstream parser that simply calls json.loads() with no
fence stripping, no repair, and no retry, exactly like a real pipeline would.
Temperature stays at 0.7 for all three strategies. The only variable is how tightly the output is constrained, so any difference in reliability comes from the constraint, not from turning down randomness.
Theory & Concepts#
1. A correct answer can still be unusable
Unconstrained, the model usually understands the review perfectly — and answers in prose. To a program that expects JSON, that is a crash on every run. The model is not wrong; it is unusable as a software component.
2. Persuasion vs. structural guarantee
- A · Unconstrained: high variance, near-0% parseable.
- B · Prompt asks for JSON: most runs parse, but anything short of 100% is a silent failure rate you must wrap in retry and repair code (markdown fences, extra prose).
- C · Schema enforced:
response_formatwith a strict JSON schema constrains decoding itself, so every run parses and matches the shape.
3. Metrics reported
- Unique responses / consistency: how many distinct strings came back, and the share taken by the most common one.
- Parses as JSON / matches schema: what a downstream consumer actually experiences. API failures are excluded so infrastructure errors are not mistaken for model behaviour.
Request flow#
Code flow#
review] -->|POST /variance| B[app.py
variance_route] B -->|review| C[variance.py
run_variance] C -->|review| P[build_prompts] P -->|plain + JSON prompts| C C -->|15 jobs| D[config.py
parallel_map] D -->|prompt, format| E[call_model] E -->|chat.completions| F[OpenAI API
gpt-4o-mini] F -->|text| E E -->|15 responses| C C -->|responses| G[metrics
parse_downstream] G -->|unique, parse%, schema%| C C -->|text report| B B -->|JSON result| A
Source Code#
Every Python file this demo runs, complete and unedited: the feature code first, then the shared Flask routes.
variance.py#
Variance & Determinism: how output constraints make a model a reliable component.
"""Variance & Determinism: how output constraints make a model a reliable component.
Adapted from ``study/09-ai-reliability/variance-determinism.py``. The same
extraction task runs N times under three strategies at temperature 0.7:
A Unconstrained prompt
B Prompt asks for JSON
C Schema enforced by the API (``response_format=json_schema``, strict)
Each response is graded by a strict downstream parser (plain ``json.loads``,
no fence stripping or repair), so the report shows how often a real pipeline
would survive the output.
"""
import json
from collections import Counter
from config import CHAT_MODEL, get_openai_client, parallel_map
RUNS_PER_STRATEGY = 5
TEMPERATURE = 0.7
MAX_REVIEW_CHARS = 1000
STRATEGY_CHOICES = ("all", "A", "B", "C")
DEFAULT_REVIEW = (
"The new smartphone is amazing, the camera quality is top-notch but the "
"battery life is a bit disappointing. I love the design though!"
)
RESPONSE_SCHEMA = {
"type": "object",
"properties": {
"sentiment": {"type": "string", "enum": ["Positive", "Negative", "Mixed"]},
"entities": {"type": "array", "items": {"type": "string"}},
},
"required": ["sentiment", "entities"],
"additionalProperties": False,
}
SCHEMA_FORMAT = {
"type": "json_schema",
"json_schema": {"name": "review_extraction", "strict": True, "schema": RESPONSE_SCHEMA},
}
def build_prompts(review: str) -> tuple[str, str]:
"""Return (unconstrained_prompt, json_prompt) for a review."""
# ① describe the extraction task with no output-shape guarantee
task = f"Extract the sentiment and key entities from this customer review: '{review}'"
# ② add plain-language JSON instructions for the prompt-only strategy
constrained = (
f"{task}\n"
"Output only a JSON object with the keys 'sentiment' and 'entities'. "
"'sentiment' should be a string (Positive, Negative, or Mixed). "
"'entities' should be a list of strings."
)
return task, constrained
def call_model(prompt: str, response_format: dict | None = None, temperature: float = TEMPERATURE) -> str:
"""Return the model's text, or a string starting with 'Error:' on failure."""
# ① build the shared chat-completion arguments for this strategy
kwargs = {
"model": CHAT_MODEL,
"messages": [{"role": "user", "content": prompt}],
"temperature": temperature,
}
if response_format is not None:
# ② attach the schema contract only for the API-enforced strategy
kwargs["response_format"] = response_format
try:
# ③ call the model and return the raw text the downstream parser will see
response = get_openai_client().chat.completions.create(**kwargs)
return (response.choices[0].message.content or "").strip()
except Exception as e:
# ④ capture provider failures so the report can exclude them explicitly
return f"Error: {e}"
def parse_downstream(raw: str) -> tuple[bool, bool]:
"""Return (parsed_ok, schema_ok) the way a strict pipeline would see it."""
# ① reject captured provider failures before attempting JSON parsing
if raw.startswith("Error:"):
return False, False
try:
# ② parse exactly what the model emitted, with no cleanup or repair
data = json.loads(raw)
except (json.JSONDecodeError, ValueError):
return False, False
# ③ require an object before checking the expected fields
if not isinstance(data, dict):
return True, False
sentiment = data.get("sentiment")
entities = data.get("entities")
# ④ validate the minimal schema the downstream application needs
schema_ok = (
sentiment in ("Positive", "Negative", "Mixed")
and isinstance(entities, list)
and all(isinstance(e, str) for e in entities)
)
return True, schema_ok
def metrics(responses: list[str]) -> dict:
"""Variance and downstream-reliability metrics; API errors are excluded."""
# ① separate provider errors from responses a downstream parser could process
valid = [r for r in responses if not r.startswith("Error:")]
errors = len(responses) - len(valid)
if not valid:
return {"unique": 0, "consistency": 0, "parse": 0, "schema": 0, "total": 0, "errors": errors}
# ② score each valid response for JSON parsing and schema compliance
verdicts = [parse_downstream(r) for r in valid]
n = len(valid)
# ③ summarize variance, consistency, and reliability percentages
return {
"unique": len(Counter(valid)),
"consistency": round(Counter(valid).most_common(1)[0][1] / n * 100),
"parse": round(sum(p for p, _ in verdicts) / n * 100),
"schema": round(sum(s for _, s in verdicts) / n * 100),
"total": n,
"errors": errors,
}
def run_variance(review: str, strategy: str = "all", temperature: float = TEMPERATURE) -> str:
"""Run the selected strategy (or all three) on ``review`` and return a text report."""
# ① trim learner input to a safe demo length and prepare both prompt variants
review = review.strip()[:MAX_REVIEW_CHARS]
unconstrained, constrained = build_prompts(review)
# ② define the three reliability strategies from loosest to strictest
strategies = [
("A", "A · Unconstrained prompt", unconstrained, None),
("B", "B · Prompt asks for JSON", constrained, None),
("C", "C · Schema enforced by API", constrained, SCHEMA_FORMAT),
]
if strategy != "all":
# ③ keep only the requested strategy when the UI filter is used
strategies = [s for s in strategies if s[0] == strategy]
# ④ run repeated calls for every selected strategy in parallel
jobs = [(p, fmt, temperature) for _, _, p, fmt in strategies for _ in range(RUNS_PER_STRATEGY)]
outputs = parallel_map(lambda job: call_model(*job), jobs)
# ⑤ build a report showing parser survival and a sample output per strategy
lines = [
f"{RUNS_PER_STRATEGY} runs per strategy · {CHAT_MODEL} · temperature {temperature}",
"",
]
schema_scores = []
for i, (_, label, _, _) in enumerate(strategies):
# ⑥ compute metrics for this strategy's batch and append them to the report
responses = outputs[i * RUNS_PER_STRATEGY:(i + 1) * RUNS_PER_STRATEGY]
m = metrics(responses)
schema_scores.append(m["schema"])
lines.append(label)
lines.append(
f" unique responses {m['unique']}/{m['total']} · consistency {m['consistency']}% · "
f"parses as JSON {m['parse']}% · matches schema {m['schema']}%"
)
if m["errors"]:
lines.append(f" ({m['errors']} API call(s) failed and were excluded)")
sample = next((r for r in responses if not r.startswith("Error:")), responses[0])
lines.append(f" sample: {sample[:300]}{'…' if len(sample) > 300 else ''}")
lines.append("")
# ⑦ finish with the full comparison takeaway or a tip for filtered runs
if len(schema_scores) == 3:
lines.append(
f"Takeaway: usable output went {schema_scores[0]}% → {schema_scores[1]}% → {schema_scores[2]}% "
"at the same temperature. Determinism came from narrowing what the model "
"was allowed to emit, not from turning down randomness."
)
else:
lines.append("Tip: re-run at another temperature, or pick 'All three' to compare strategies.")
return "\n".join(lines)
if __name__ == "__main__":
# ① run the default review when this module is executed directly
print(run_variance(DEFAULT_REVIEW))
config.py#
Shared configuration: load .env and build the OpenAI client.
"""Shared configuration: load .env and build the OpenAI client.
This module is the single place that knows about secrets and model names.
Every feature module imports from here instead of reading ``os.environ`` or
constructing API clients itself.
"""
import os
from concurrent.futures import ThreadPoolExecutor
from dotenv import load_dotenv
from openai import OpenAI
load_dotenv()
def get_env(name: str, default: str = "") -> str:
"""Return an environment variable, falling back to ``default``."""
return os.environ.get(name, default)
# A NON-reasoning model: reasoning models think internally even when told not
# to, which erases the gaps these demos measure.
CHAT_MODEL = get_env("OPENAI_MODEL", "gpt-4o-mini")
# Upper bound on parallel API calls per request; keeps each demo well under
# the 60 s Nginx proxy timeout without hammering the provider.
MAX_WORKERS = 8
# Temperatures selectable from the UI, keyed by the string the browser sends.
TEMPERATURE_CHOICES = {"0": 0.0, "0.7": 0.7, "1.2": 1.2}
_client = None
def get_openai_client() -> OpenAI:
"""Return a shared OpenAI client built from OPENAI_API_KEY."""
global _client
if _client is None:
# ① read the API key lazily so tests/imports do not require credentials
api_key = get_env("OPENAI_API_KEY")
if not api_key:
raise RuntimeError("OPENAI_API_KEY is not set. Add it to your .env file.")
# ② create one reusable client with bounded timeout and retries
_client = OpenAI(api_key=api_key, timeout=30, max_retries=2)
return _client
def parallel_map(fn, items):
"""Run ``fn`` over ``items`` concurrently, preserving input order."""
with ThreadPoolExecutor(max_workers=MAX_WORKERS) as pool:
return list(pool.map(fn, items))
def bar(correct: int, total: int, width: int = 10) -> str:
"""Render a text accuracy bar with a count and percentage."""
# ① avoid dividing by zero when a filtered comparison has no examples
if total == 0:
return "n/a"
# ② convert the score into filled and empty bar characters
filled = round(correct / total * width)
return f"{'█' * filled}{'░' * (width - filled)} {correct}/{total} ({round(correct / total * 100)}%)"
app.py#
Flask server for AI Reliability Lab: four measured LLM reliability demos.
"""Flask server for AI Reliability Lab: four measured LLM reliability demos.
Architecture notes
------------------
- All routes are attached to a Blueprint (``bp``) instead of directly to
``app``. This lets us register the entire Blueprint under a runtime URL
prefix (``PATH_PREFIX``) without touching individual route strings.
- In local development PATH_PREFIX is empty, so routes are at "/",
"/variance", etc. In production Nginx forwards ``/ai-reliability/...``
traffic to the container and PATH_PREFIX is set to "/ai-reliability".
- flask-cors adds ``Access-Control-Allow-Origin: *`` headers so the HTML
page can call the API even if it is served from a different origin during
development.
"""
import os
from pathlib import Path
from flask import Blueprint, Flask, jsonify, request
from flask_cors import CORS
from config import TEMPERATURE_CHOICES
from cot import PROBLEMS, run_cot
from cot import STRATEGY_CHOICES as COT_STRATEGIES
from rate_limiter import check_rate_limit
from tool_errors import FAIL_ON_CALL_LABELS, PROMPT_CHOICES, run_error_injection
from tool_routing import DESCRIPTION_CHOICES, NAME_CHOICES, run_routing
from variance import STRATEGY_CHOICES as VARIANCE_STRATEGIES
from variance import run_variance
# ---------------------------------------------------------------------------
# Configuration
# ---------------------------------------------------------------------------
# PATH_PREFIX is set by the deployment environment ("/ai-reliability") so
# the app works correctly behind an Nginx location block. Locally it is an
# empty string, which mounts all routes at the root.
PATH_PREFIX = os.environ.get("PATH_PREFIX", "")
# app.py lives in src/python, while index.html, css/, and js/ live in src/.
STATIC_DIR = Path(__file__).resolve().parents[1]
app = Flask(__name__, static_folder=str(STATIC_DIR))
# Allow cross-origin requests from any origin. In production you would
# restrict this to the specific front-end domain.
CORS(app)
# A Blueprint groups related routes. We register it once at the bottom with
# the runtime PATH_PREFIX, avoiding any hardcoded path strings in the routes.
bp = Blueprint("main", __name__)
@bp.before_request
def enforce_rate_limit():
"""Enforce strict 10 requests per hour limit on all POST endpoints."""
# ① apply the limit only to API actions, not static page loads
if request.method == "POST":
# ② ask the shared limiter whether this request should be blocked
blocked, msg, retry_after = check_rate_limit(
request, max_requests=10, window_seconds=3600
)
if blocked:
# ③ return a 429 with retry guidance when the hourly quota is exhausted
resp = jsonify({"error": msg})
resp.status_code = 429
resp.headers["Retry-After"] = str(retry_after)
return resp
# ---------------------------------------------------------------------------
# Routes
# ---------------------------------------------------------------------------
@bp.route("/")
def index():
"""Serve index.html, injecting the correct API base URL for the environment."""
# ① read the static HTML shell from the configured Flask static folder
with open(os.path.join(app.static_folder, "index.html"), encoding="utf-8") as f:
html = f.read()
# ② inject the runtime path prefix so browser fetches target the right API base
# The HTML file ships with 'data-api-base=""' (empty = relative URL, works
# locally). For production we replace it with the actual path prefix so
# all fetch() calls in the browser target the right endpoint.
html = html.replace('data-api-base=""', f'data-api-base="{PATH_PREFIX}"')
# ③ return the modified HTML with an explicit text/html response type
return app.response_class(html, mimetype="text/html")
@bp.route("/css/<path:filename>")
def css(filename):
"""Serve stylesheets from the src/css directory."""
return app.send_static_file(os.path.join("css", filename))
@bp.route("/js/<path:filename>")
def js(filename):
"""Serve scripts from the src/js directory."""
return app.send_static_file(os.path.join("js", filename))
@bp.route("/info/<path:filename>")
def info(filename):
"""Serve the "how this demo works" explainer pages from src/info."""
return app.send_static_file(os.path.join("info", filename))
def read_message() -> str:
"""Return the trimmed ``message`` field from the JSON body, or ''."""
# ① parse JSON leniently so missing or malformed bodies become empty data
data = request.get_json(force=True, silent=True) or {}
# ② normalize the message field into a stripped string for route validation
return str(data.get("message") or "").strip()
def read_choice(name: str, allowed, default: str) -> str | None:
"""Return a dropdown value from the JSON body, ``default`` if absent, or None if not allowed."""
# ① parse JSON leniently and fall back to the route's default choice
data = request.get_json(force=True, silent=True) or {}
value = str(data.get(name) or default)
# ② accept only known UI choices so feature modules receive valid selectors
return value if value in allowed else None
def invalid_choice(name: str, allowed):
return jsonify({"error": f"Invalid {name}. Choose one of: {', '.join(map(str, allowed))}."}), 400
@bp.route("/variance", methods=["POST"])
def variance_route():
"""Run Unconstrained vs Prompt-JSON vs Schema-enforced extraction on a review."""
# ① validate that the learner supplied review text to analyze
message = read_message()
if not message:
return jsonify({"error": "A customer review is required."}), 400
# ② read and validate the selected extraction strategy
strategy = read_choice("strategy", VARIANCE_STRATEGIES, "all")
if strategy is None:
return invalid_choice("strategy", VARIANCE_STRATEGIES)
# ③ read and validate the selected temperature
temperature = read_choice("temperature", TEMPERATURE_CHOICES, "0.7")
if temperature is None:
return invalid_choice("temperature", TEMPERATURE_CHOICES)
try:
# ④ run the variance feature and return its text report as JSON
return jsonify({"result": run_variance(message, strategy, TEMPERATURE_CHOICES[temperature])})
except Exception:
# ⑤ log server-side detail while returning a safe client-facing error
app.logger.exception("variance failed")
return jsonify({"error": "Variance test failed. Please try again later."}), 500
@bp.route("/cot", methods=["POST"])
def cot_route():
"""Run Direct vs Chain-of-Thought prompting on one preset problem."""
# ① read the requested problem key and normalize it for lookup
message = read_message().lower()
if not message:
return jsonify({"error": "A problem name is required."}), 400
if message not in PROBLEMS:
return jsonify({"error": f"Unknown problem. Choose one of: {', '.join(PROBLEMS)}."}), 400
# ② read and validate the selected prompt strategy
strategy = read_choice("strategy", COT_STRATEGIES, "both")
if strategy is None:
return invalid_choice("strategy", COT_STRATEGIES)
# ③ read and validate the selected temperature
temperature = read_choice("temperature", TEMPERATURE_CHOICES, "0.7")
if temperature is None:
return invalid_choice("temperature", TEMPERATURE_CHOICES)
try:
# ④ run the CoT feature and return its text report as JSON
return jsonify({"result": run_cot(message, strategy, TEMPERATURE_CHOICES[temperature])})
except Exception:
# ⑤ log server-side detail while returning a safe client-facing error
app.logger.exception("cot failed")
return jsonify({"error": "CoT comparison failed. Please try again later."}), 500
@bp.route("/routing", methods=["POST"])
def routing_route():
"""Route a weather question under the names x descriptions 2x2."""
# ① validate that the learner supplied a weather-routing question
message = read_message()
if not message:
return jsonify({"error": "A weather question is required."}), 400
# ② read and validate the selected tool-name condition
names = read_choice("names", NAME_CHOICES, "both")
if names is None:
return invalid_choice("names", NAME_CHOICES)
# ③ read and validate the selected tool-description condition
descriptions = read_choice("descriptions", DESCRIPTION_CHOICES, "both")
if descriptions is None:
return invalid_choice("descriptions", DESCRIPTION_CHOICES)
try:
# ④ run the routing feature and return its text report as JSON
return jsonify({"result": run_routing(message, names, descriptions)})
except Exception:
# ⑤ log server-side detail while returning a safe client-facing error
app.logger.exception("routing failed")
return jsonify({"error": "Routing test failed. Please try again later."}), 500
@bp.route("/errors", methods=["POST"])
def errors_route():
"""Run the tool-error injection scenarios on a weather question."""
# ① validate that the learner supplied a weather question for the agent
message = read_message()
if not message:
return jsonify({"error": "A weather question is required."}), 400
# ② read and validate which system-prompt scenario to run
prompt = read_choice("prompt", PROMPT_CHOICES, "both")
if prompt is None:
return invalid_choice("prompt", PROMPT_CHOICES)
# ③ derive valid failure-injection choices from the shared labels
fail_choices = [str(k) for k in FAIL_ON_CALL_LABELS]
fail_on_call = read_choice("fail_on_call", fail_choices, "2")
if fail_on_call is None:
return invalid_choice("fail_on_call", fail_choices)
try:
# ④ run the error-injection feature and return its text report as JSON
return jsonify({"result": run_error_injection(message, prompt, int(fail_on_call))})
except Exception:
# ⑤ log server-side detail while returning a safe client-facing error
app.logger.exception("error injection failed")
return jsonify({"error": "Error-injection test failed. Please try again later."}), 500
# ---------------------------------------------------------------------------
# Blueprint registration + server entry point
# ---------------------------------------------------------------------------
# Register all Blueprint routes under the optional path prefix. This single
# line is the only place where PATH_PREFIX is applied — every route above is
# written as a relative path (e.g. "/variance") and the prefix is prepended here.
app.register_blueprint(bp, url_prefix=PATH_PREFIX)
if __name__ == "__main__":
# Run the development server. 0.0.0.0 makes the app reachable from outside
# the container; port 5000 is mapped to the host port in docker-compose.yml.
app.run(host="0.0.0.0", port=5000)