RESEARCH · EVALS ENGINEERING
Case studies in LLM evaluation
These are engineering postmortems, not case studies in the marketing sense. Each one documents what broke, what was measured, and what still doesn't work. The autopsy section is mandatory and never softened.
- CS-01→When my LLM scorer told me what I wanted to hearRecalibrating a job-fit pipeline from 93% false confidence to deduction-based honesty.LLM EvaluationPipeline CalibrationAgentic Systems
- CS-02Designing tasks that break frontier modelsAn adversarial benchmark methodology for capability elicitation at Terminal Bench.BenchmarkingAdversarial EvalsFrontier ModelsIN PROGRESS
- CS-03Skills as prompt architectureA multi-agent creative system with a PASS/FAIL critic loop — plasmar + alma-critic.Multi-agentPrompt ArchitectureCritic SystemsIN PROGRESS
CS-02 and CS-03 are in progress — content published once the work is complete enough to be honest about what failed.