← AG.
RESEARCH · EVALS ENGINEERING

Case studies in LLM evaluation

These are engineering postmortems, not case studies in the marketing sense. Each one documents what broke, what was measured, and what still doesn't work. The autopsy section is mandatory and never softened.

  1. CS-01
    When my LLM scorer told me what I wanted to hear
    Recalibrating a job-fit pipeline from 93% false confidence to deduction-based honesty.
    LLM EvaluationPipeline CalibrationAgentic Systems
  2. CS-02
    Designing tasks that break frontier models
    An adversarial benchmark methodology for capability elicitation at Terminal Bench.
    BenchmarkingAdversarial EvalsFrontier ModelsIN PROGRESS
  3. CS-03
    Skills as prompt architecture
    A multi-agent creative system with a PASS/FAIL critic loop — plasmar + alma-critic.
    Multi-agentPrompt ArchitectureCritic SystemsIN PROGRESS

CS-02 and CS-03 are in progress — content published once the work is complete enough to be honest about what failed.