AI Systems Lab
Prompting trade-offs, model sycophancy, and multi-judge agentic systems
⚖️ Prompting Benchmark ⓘ
Compare Direct, Zero-Shot CoT, and Few-Shot CoT on multi-step reasoning questions. Pick a preset (scored against its known answer) or choose Custom and type your own.
🛡️ Sycophancy Trap ⓘ
A cold question followed by up to three rounds of escalating user pressure (pushback, refusal, compliance demand). Pick a case and how hard to push.
🏛️ The Refund Bench ⓘ
Four-stage agentic dispute resolution pipeline: freezes evidence, extracts atomic grievances with co-reference resolution, runs a bench of judge models with majority voting, and settles with code pricing under an auto-approval cap. Change the bench or the cap and re-run to see how rulings and payouts shift.