The Story That Started It All

It was a Tuesday. One of those Tuesdays where you think you’re going to change the world.

I had just pushed my first real LLM feature to production: a customer support bot for an internal tool. It was beautiful. It used GPT-4, it had a system prompt that I spent 3 days writing, and I was convinced it was going to make our users weep with joy.

Then my client pings me:

“Dhanush, why are we spending $200/month on evals for a bot that answers 12 questions a day?”

I had set up an eval pipeline for everything. Every. Single. Output. Classification evals. Coherence evals. Tone evals. Factuality evals. I even had an eval checking if the model was “being polite enough.”

The model was checking if the model was being polite. I was paying for a robot to judge another robot on manners. My parents did not fund my JEE coaching for this.

That week, I learned something that nobody in the AI hype train talks about:

Evals are a tool. Not a religion.

Let’s talk about when you actually need them and when you’re just LARPing as a serious ML engineer.

First, What Even Are Evals?

Okay, quick 30-second explainer before we go deeper.

Evals = Evaluations. It’s how you measure whether your LLM is doing what you want it to do.

There are two main types:

  • Human Evals: A human reads outputs and rates them. Cost: Expensive (time)
  • LLM-as-Judge Evals: Another model judges your model’s output Cost: Expensive (tokens + $$$)
  • Heuristic/Code Evals: Regex, JSON checks, unit tests Cost: Basically free (btw, I’m huge fan of this)

Most people default to LLM-as-judge for everything because it feels fancy and “AI-native.”

That’s the mistake. Let’s fix it.

The Big Mental Model: A Decision Tree

Before you run any eval, ask yourself this:

Evals Usage Decision Flow — Dhanush Kandhan

Bookmark this. Tattoo it. Make it your phone wallpaper. I don’t care. Just internalize it.

Chapter 1: When You Absolutely DO Need Evals

1. You’re Changing the Model or Prompt and Need Regression Testing

You’ve been running your system on GPT-4o. Your boss/manager says switch to Claude 3.5 to save costs. Or you’re tweaking your system prompt and you don’t want to break what’s already working.

This is eval territory. You need a golden dataset of inputs + expected outputs (or qualities), and you want to make sure your changes don’t make things worse.

The rule of thumb: If you’re shipping a change that affects model behavior, you need evals before that change goes to users.

No evals here = you’re testing in production = you’re a cowboy = respect, but also please don’t.

2. The Output Quality Is Subjective and High Stakes

You’re building a medical summarization tool. Or a legal document reviewer. Or an AI tutor for kids.

These are cases where “it works most of the time” isn’t enough. A subtle factual error in a medical context isn’t just bad UX: it’s a lawsuit, or worse.

Here, LLM-as-judge evals with a strong rubric make sense. You want to catch when the model hallucinates, skips critical information, or makes confident-sounding but wrong claims.

The cost is justified because the cost of getting it wrong is massive.

3. You’re Building a Feedback Loop / Fine-Tuning Pipeline

If you’re collecting model outputs to eventually fine-tune or RLHF your model, you need evals to filter the good data from the bad.

Simple! Garbage in = garbage out.

Especially true when you’re training on your own outputs. (This is how model collapse happens, and no, it’s not as dramatic as it sounds but it’s very real.)

Evals here act as a quality gate for your training data. Non-negotiable.

4. A/B Testing Prompts at Scale

You’ve written two versions of a system prompt. Version A is concise. Version B adds a few reasoning steps. You want to know which one performs better across 500 diverse inputs not just the 5 examples you cherry-picked. Btw, I built a tool for this to make prompt management (atomspect.letretro.com) still in pre-alpha with real users.

This is exactly what evals are built for. Run both prompts against your eval set. Compare scores. Ship the winner.

Without this, you’re just guessing and “gut feeling” is not a valid ML metric, even if your gut went to IIT.

Chapter 2: When You Absolutely DO NOT Need Evals

This is the part nobody writes about because it doesn’t sound impressive at conferences.

1. Your Output Is Structured and Parseable

You asked the model to return JSON. Either it returns valid JSON or it doesn’t.

import json

def check_output(response):
try:
parsed = json.loads(response)
assert "title" in parsed
assert "summary" in parsed
return True
except:
return False

Done. That’s your eval. It’s a unit test. It runs in 2 milliseconds and costs literally nothing. You do NOT need GPT-4 to judge if the JSON has the right keys.

I’ve seen people spend $200/month having an LLM check if another LLM’s JSON output was valid JSON.

The JSON parser is free. Use the JSON parser.

2. You Haven’t Done Basic Prompt Engineering Yet

Here’s an embarrassing truth from my own past: I once ran 2,000 LLM-as-judge evals to figure out why my model kept generating outputs that were “too formal.”

Know what fixed it?

Adding "Write in a casual, friendly tone" to my system prompt.

Evals don’t fix bad prompts. They just expensively confirm that your prompt is bad.

The order of operations should always be:

  1. Write prompt
  2. Test on 10–20 examples manually
  3. Iterate prompt
  4. When prompt feels stable → run evals to confirm

Don’t jump to step 4 while you’re still figuring out step 1. It’s like running a marathon before you’ve learned to walk. Very embarrassing. Also expensive.

3. The Task is Simple and Low Stakes

You’re using an LLM to reformat dates from “May 30, 2026” to “2026–05–30.”

You don’t need an eval pipeline. You need to check the output format. With code. That runs in microseconds.

Apply the “would I write a unit test for this in normal software engineering?” filter. If yes → write the unit test, not an eval.

4. You’re in Early Exploration Mode

You’re prototyping. You’re figuring out if the idea even works. You’ve written the first version of your prompt last Thursday.

Evals at this stage are a trap. Your prompt will change so dramatically in the next week that any eval dataset you build today will be irrelevant by Monday.

Explore first. Stabilize. Then evaluate.

Think of it like buying furniture for an apartment you haven’t designed yet. You’re going to move it anyway. Maybe do the floor plan first.

Chapter 3: The “Is This a Prompt Problem or a Model Problem?” Question

This one trips up a lot of people, including me, especially me. (naanae dan)

Here’s the rule:

Before you eval, diagnose.

Problem Analysis Table while working on Evals/ Prompt Engineering — Dhanush Kandhan

Only that last row is an “eval first” situation. Everything above it is a prompt engineering problem dressed up as a model problem.

Chapter 4: What to Put IN Your Evals (When You Do Need Them)

Okay, so you’ve gone through the decision tree. You’ve done your prompt engineering. You’ve confirmed this is actually an eval situation. Now what?

The Golden Dataset Is Everything

Your eval is only as good as your test cases. Here’s how to build a solid one:

Include edge cases, not just happy paths.

Your eval set should have:

  • The easy cases (duh)
  • The weird edge cases (what happens when input is in Tamil? what if someone sends emoji? what if the query is ambiguous?)
  • The adversarial cases (what if someone tries to jailbreak? what if the input is deliberately confusing?)

Aim for 50–200 examples to start. Not 10 (too small, noisy). Not 2000 (too expensive for iteration). 50–200 is the sweet spot for most use cases in early production.

Label your expected behavior, not expected output. Don’t write evals like “the output should be exactly: ‘Your order has been placed.’” Write evals like “the output should confirm the order, mention the order ID, and not include any pricing information.”

The first is brittle. The second is robust.

Pick the Right Judge for the Right Job

  1. Eval Method JSON/schema validation — Code (free, fast)
  2. Classification accuracy — Code comparison (free)
  3. Factual correctness (with source) — Code/heuristic + string match
  4. Tone, style, helpfulness — LLM-as-judge (worth the cost)
  5. Safety/harmful content — Specialized classifiers or LLM-as-judge
  6. Coherence/fluency — LLM-as-judge or human spot-check

Don’t use a Ferrari to go to the grocery store. (Unless you’re a different kind of FAANG engineer, in which case, please pay for my samosas.)

Chapter 5: The Real Cost of Blind Evals

Let me give you numbers, because engineers love numbers and also because I lived this.

Say you’re running evals on 1,000 outputs per day using GPT-4o as judge.

  • Average tokens per eval call: ~1,500 (input prompt + output being judged + judge prompt)
  • 1,000 evals/day = 1.5M tokens/day
  • At roughly $5/1M input tokens for GPT-4o: ~$7.50/day just for evals
  • Per month: ~$225

Now ask yourself: is your eval actually catching anything? Is it changing decisions you make? Are you reading these eval results?

If the answer to any of those is “not really”, you’re lighting $225/month on fire to feel rigorous.

That’s 3 months of a good Spotify + Netflix combo. That’s a nice mechanical keyboard. That’s 180 cups of good coffee at Blue Tokai.

Spend it better.

The Final Checklist: Before You Set Up That Eval Pipeline

Print this. Stick it above your monitor. Especially if you’re working at 2 AM on a startup idea:

  • Have I manually tested this with 10–20 diverse examples?
  • Is my prompt stable, or am I still changing it every day?
  • Can this be checked with code instead of an LLM?
  • Is this task high-stakes enough to justify the cost?
  • Do I have a golden dataset that covers edge cases?
  • Will I actually act on these eval results?
  • Am I using the cheapest judge that can do the job?
  • Is this a prompt problem I haven’t fixed yet?

If you can check all 8 boxes: run your evals. You’ve earned it.

If you can’t: go fix the thing that’s not checked first.

TL;DR (For My Fellow Engineers Who Read Bottom-Up)

If your output is structured and parseable: JSON, schema, a field you can check, write a unit test and move on. Same goes for simple, low-stakes tasks: a quick manual spot check is all you need, no LLM judge required.

And if you’re still actively changing your prompt every other day, stop optimize the prompt first, then think about evals. Running evals on an unstable prompt is like benchmarking a car you haven’t finished building yet.

Now, when do you pull out the evals? When you’re changing a model or prompt in production and need regression coverage, that’s non-negotiable. When your output quality is subjective and the stakes are high (medical, legal, anything where wrong = bad), LLM-as-judge evals earn their cost. When you’re building a fine-tuning pipeline, evals are your quality gate between good training data and garbage.

And when you’re A/B testing prompts at scale and need real signal instead of vibes, run the evals, compare, ship the winner.

Closing Thoughts (From Someone Who’s Been There)

Look, evals are genuinely powerful. They’re the difference between “I think this works” and “I know this works, here’s the data.” That confidence is worth a lot when you’re shipping AI features to real users.

But evals aren’t magic. They don’t fix broken prompts. They don’t justify skipping engineering fundamentals. And they definitely don’t need to run on every single output of every single feature you build.

Be intentional. Be cheap where you can. Be rigorous where it matters.

And please, for the love of everything good, do not use GPT-4 to check if your JSON is valid JSON. I’m begging you.

Now go ship something. Validate it properly. And maybe get some sleep.

Found this useful? Share it with that one teammate who sets up a 500-eval pipeline on day 1 of every project. They need this more than they know.

Thank You :) See you in next!!!