AI Reliability Lab

Measure the levers that turn a flaky model into a dependable component

🎲 Variance & Determinism ⓘ

Extracts sentiment and entities from a review 5 times per strategy (unconstrained, prompt asks for JSON, API-enforced schema), then grades every output with a strict json.loads() parser. Try temperature 0 vs 1.2 to see what randomness does and does not fix.

🧠 Chain-of-Thought vs Direct ⓘ

Runs a multi-step reasoning problem 5 times per strategy — a direct prompt and "think step by step" — scoring only the final line.

🧭 Tool Schema Calibration ⓘ

Offers three weather tools (current, forecast, history) under a 2x2 of descriptive vs opaque names and loose vs tight descriptions. Your question is routed under each selected condition, then 12 labelled queries score its accuracy.

🚨 Tool Error Injection ⓘ

An agent answers with a get_weather tool that returns HTTP 503 on the chosen call. Compare a plain prompt with one that demands error reporting: did the model admit the failure, give the code, or invent weather? Ask about two or more cities.