AI Reliability Lab
Measure the levers that turn a flaky model into a dependable component
🎲 Variance & Determinism ⓘ
Extracts sentiment and entities from a review 5 times per strategy (unconstrained, prompt
asks for JSON, API-enforced schema), then grades every output with a strict
json.loads() parser. Try temperature 0 vs 1.2 to see what randomness does
and does not fix.
🧠 Chain-of-Thought vs Direct ⓘ
Runs a multi-step reasoning problem 5 times per strategy — a direct prompt and "think step by step" — scoring only the final line.
🧭 Tool Schema Calibration ⓘ
Offers three weather tools (current, forecast, history) under a 2x2 of descriptive vs opaque names and loose vs tight descriptions. Your question is routed under each selected condition, then 12 labelled queries score its accuracy.
🚨 Tool Error Injection ⓘ
An agent answers with a get_weather tool that returns HTTP 503 on the chosen
call. Compare a plain prompt with one that demands error reporting: did the model admit the
failure, give the code, or invent weather? Ask about two or more cities.