This post summarizes the principles I keep in mind whenever I use LLM-as-a-Judge to evaluate and improve a system. Most of them come from my experience building LLM Judges while working at Hyperconnect1 and from Hamel Husain’s writing.2

What I want to discuss here is not how to prompt an LLM Judge. It is how to treat an LLM Judge as an ML system component and build an evaluation loop that lets us continuously optimize a system in the direction we actually want.

You Can’t Improve What You Can’t Measure

It is a well-known saying. If we cannot measure which version of a system is better, we cannot know what to improve. On the other hand, if we can measure the direction we want with a reasonably quantitative signal, we can use that signal to train models, compare experiments, analyze failures, and continuously improve the system.

A similar idea comes up frequently in AI today: if a problem is verifiable, it can eventually be optimized through reinforcement learning or a related method.3 The reward does not have to be perfect. If we can produce a signal that is sufficiently related to the direction we want, we can begin an optimization loop.

For generative features or search results whose quality is difficult to evaluate programmatically, it is natural to consider LLM-as-a-Judge. For example, a chatbot team may want to evaluate the helpfulness, conciseness, or correctness of a response. A search team may want to measure query-document relevance. In the past, humans labeled these subjective qualities directly. Now, part of that work can be delegated to an LLM Judge.

But introducing an LLM Judge immediately requires caution. Just because an LLM Judge produces a number does not mean that increasing that number will produce the system we actually want. If we optimize the Judge score aggressively but the real user experience or business outcome does not improve at all, we have not optimized the system well. We may simply have optimized the wrong objective very efficiently.

Before building an LLM Judge, it is therefore worth taking one step back.

Do You Really Need an LLM Judge?

First, ask whether there are any programmatic or directly observable signals that can evaluate the quality you care about.

Consider approaches such as RLVR, or Reinforcement Learning with Verifiable Rewards. In coding tasks, we can programmatically verify whether the final answer is correct, whether the code runs, whether it passes tests, and whether it satisfies the required format.

These signals do not perfectly express every kind of code quality we care about. Passing tests does not necessarily mean that code is readable or maintainable. In other words, these signals are incomplete with respect to the quality we ultimately want. But they are still highly reliable for an important part of that quality.

To simplify the idea, let $U$ be the user utility we ultimately want to optimize and $M$ be a measurable proxy signal. We do not need:

\[M = U\]

In practice, finding such a perfect metric is difficult.

What matters instead is whether the following relationship generally holds:

\[M \uparrow \quad \Rightarrow \quad U \text{ also tends to } \uparrow\]

The first question is whether an imperfect signal is still useful enough to move the system in the right direction.

For a consumer product, when evaluation is difficult, a useful starting point is the business objective: what does the product ultimately need to accomplish? From there, work backward through the user journey or funnel to identify measurable signals.

Suppose we are building an AI chatbot for customer support. Its ultimate business objective might be:

\[\begin{aligned} \min \quad &\text{Customer Support Cost} \\ \text{subject to} \quad &\text{Customer Satisfaction} \geq C \end{aligned}\]

In other words, the goal is to reduce support costs while keeping customer satisfaction above an acceptable level.

What does reducing support costs mean in practice?

One proxy is whether the AI chatbot resolves the user’s problem before escalating it to a human agent. We can measure the escalation rate. But a lack of escalation does not necessarily mean the problem was resolved. The user may simply have given up.

We can also ask whether the user reported that the problem was resolved after the conversation. We might examine whether the user contacts support again about the same issue within a certain period, or whether the ticket is reopened.

The problem can be decomposed in the following direction:

\[\begin{gathered} \text{Business Objective} \rightarrow \text{Product Outcome} \\ \rightarrow \text{User Behavior} \rightarrow \text{Measurable Signal} \end{gathered}\]

Each signal is clearly noisy and suboptimal. But taken together, several signals may cover a substantial portion of the original objective.

Before building an LLM Judge, ask:

Can programmatic or directly observable signals already capture enough of the objective we want to achieve?

If they can, those signals are the better starting point.

“It’s Hard to Eval” Is a Product Smell

If, after all this, we still cannot find any way to evaluate whether the product is working, it is worth asking whether something else is wrong. To borrow Hamel Husain’s phrase, “It’s Hard to Eval” Is a Product Smell.4

If the goal of a product is to satisfy customers but there is no way at all to determine whether customers are satisfied, we may not have defined clearly enough what we are trying to satisfy.

Of course, some final business outcomes appear only after several months. Causal attribution can be difficult, and signals can be extremely sparse. This is not an argument that every product has a perfect and immediate metric.

But a product ultimately solves some problem for a user. There must be a process through which the user moves from having that problem to having it resolved. If we understand the product and the user’s problem well enough, we should be able to break that process into smaller steps or define intermediate outcomes that provide some way to check whether the system is moving in the right direction.

If even that is difficult, evaluation may not be the only problem. We may need to revisit whether the product objective is clear enough, or whether the product itself is designed so that the user’s problem-solving process can be observed and verified.

So, When Should You Use an LLM Judge?

After thinking through the questions above, we may find several quantitative signals but still feel that they are too noisy or fail to cover an important part of the product objective. That is when an LLM Judge becomes worth considering. An LLM Judge can also be useful before a product has launched, when real user signals are not yet available.

A problem is a good candidate for an LLM Judge when it satisfies roughly the following conditions.

  • Measuring this quality would help improve the system.
  • The quality is subjective and difficult to calculate directly.
  • A sufficiently trained human can still judge it relatively consistently.
  • Repeating that judgment at scale would make it useful in a real optimization loop.

Consider query-document relevance in search. Relevance is difficult to calculate through simple string overlap. But in most cases, a domain expert who reads both the query and the document can judge whether the document is relevant to the query. This is a good candidate for an LLM Judge.

An LLM Judge is not a tool for solving problems where even humans do not know what good looks like. It is more appropriate for:

Problems that humans can judge, but where human judgment is difficult to scale.

What Makes a Good Judge?

What conditions should an LLM Judge satisfy once it has been built? There are at least three.

  • It must produce a quantifiable signal that can be used to compare systems and optimize them.
  • It must agree sufficiently with the judgment of the human experts best qualified to assess that quality.
  • Improving the resulting offline signal should move the real user signal or business outcome in the same direction.

This reveals two separate alignment problems. The first is the alignment of the Judge itself:

\[J(x) \approx H(x)\]

Here, $J$ is the judgment produced by the LLM Judge and $H$ is the judgment of a human expert.

But this alone is not enough. Ultimately, we also need:

\[H(x) \approx U(x)\]

Here, $U$ is actual user utility or the product objective.

Even if an LLM Judge achieves 95% agreement with human experts, we may still be measuring the wrong objective very precisely if the human experts’ criteria are not aligned with real user outcomes.

It is therefore important to separate the following two relationships:

\[\text{LLM Judge} \leftrightarrow \text{Human Expert}\]

and:

\[\text{Offline Evaluation} \leftrightarrow \text{Online Product Outcome}\]

The first is a Judge-calibration problem. The second asks whether the evaluation objective itself is correct.

Align Humans Before Aligning the Judge

People are often less aligned on subjective problems than we expect. Before building an LLM Judge, it is therefore useful to align the humans first.

Collect enough real system outputs, then ask the relevant domain experts, PMs, or even MLEs to label them independently. It is important to record not only a simple Pass or Fail, but also the reasoning or critique behind each judgment.

When the results are compared, there will probably be more disagreement than expected at first. These disagreements are the most valuable samples. They identify the current gray decision boundary of the product.

Comparing critiques and discussing why the judgments differ reveals implicit assumptions that the existing rubric did not specify.

\[\begin{gathered} \text{Human Disagreement} \rightarrow \text{Discussion} \\ \rightarrow \text{Implicit Assumption} \rightarrow \text{Rubric Update} \end{gathered}\]

The team can then evaluate the outputs again with the updated rubric. As this process repeats, disagreement gradually decreases and human judgments begin to converge.

Importantly, this process updates more than the rubric document. It also updates the team’s implicit understanding of the product. The decision boundary becomes clearer:

What counts as a good output for this product?

What is still acceptable, and where does it become unacceptable?

As a result, the team develops domain experts who can judge the relevant product quality with reasonable consistency.

There is also a useful side effect. Repeating this process naturally creates a human-labeled dataset.

Iterate and Update the Rubric

A rubric cannot be made perfect from the beginning. It is more natural for the rubric to emerge from real outputs and failures.

Evaluate outputs with the initial rubric, find disagreements, compare critiques, and update the rubric. Then evaluate new samples again.

\[\begin{gathered} R_0 \rightarrow \text{Label} \rightarrow \text{Disagreement} \\ \downarrow \\ R_1 \rightarrow \text{Label} \rightarrow \text{Disagreement} \\ \downarrow \\ R_2 \rightarrow \cdots \end{gathered}\]

After enough iterations, major changes to the rubric become less frequent and human agreement stabilizes. Only then can we say that “what our team considers a good output” has become reasonably explicit.

The dataset accumulated by this point can later serve as a calibration dataset for evaluating and improving the LLM Judge.

Make Evaluation as Simple as Possible

To make human alignment easier, the evaluation itself should be kept as simple as possible.

When possible, binary evaluation is a useful starting point. Using a Likert scale from 1 to 5 requires answers to questions such as:

What exactly is the difference between a 3 and a 4?

Is a 4 good enough to ship to production?

If one person gives the same output a 3 and another gives it a 4, is that truly a disagreement?

Defining the meaning of every tick becomes another difficult problem.

Binary evaluation, in contrast, focuses on a single decision boundary:

Does this output satisfy the standard we want?

Or, in more product-oriented language:

Would you ship this output?

Even if each sample is evaluated in binary terms, aggregating those decisions produces a quantitative metric:

\[\text{Pass Rate} = \frac{\#\text{Pass}}{\#\text{Total Samples}}\]

This metric can compare model A with model B or detect regressions.

Of course, binary evaluation compresses a great deal of information. The tradeoff can be framed this way:

  • Binary for optimization
  • Critiques for diagnosis

Pass or Fail supports quantitative comparison and optimization. Detailed critiques explain why the system failed and help identify the next improvement.

Treat the Judge as a Model, Not an Oracle

Once the rubric has stabilized and a human-labeled dataset exists, we can build the LLM Judge.

A reasonable starting point is to place the rubric in the prompt and provide examples and critiques written by human experts as few-shot examples.

But the important point is not to trust the LLM Judge as an oracle. The Judge is still a model. We can represent the human-labeled dataset as:

\[D_H = \{(x_i, y_i, c_i)\}_{i=1}^{N}\]

Here, $x_i$ is the output being evaluated, $y_i$ is the human Pass or Fail label, and $c_i$ is the human critique.

For the same input, the LLM Judge produces:

\[J(x_i) = (\hat{y}_i, \hat{c}_i)\]

We can begin by comparing:

\[\hat{y}_i \stackrel{?}{=} y_i\]

This measures agreement between the Judge and the human.

But aggregate agreement is not enough. We need to inspect the samples that humans labeled Fail but the Judge labeled Pass, as well as those that humans labeled Pass but the Judge labeled Fail. Then we compare the Judge’s reasoning with the human critique. Analyzing these disagreements reveals where the Judge misunderstood the rubric, missed a condition, or lacked sufficient examples.

The interesting part is that this process has almost the same structure as the one above. First, we analyzed disagreements between:

\[\text{Human} \leftrightarrow \text{Human}\]

to improve the rubric.

Now, we analyze disagreements between:

\[\text{Human} \leftrightarrow \text{Judge}\]

to improve the prompt and the Judge.

In short:

Human-human disagreement $\rightarrow$ rubric iteration

Human-Judge disagreement $\rightarrow$ Judge iteration

By repeating this process, the Judge can gradually align with the decision boundary the team has established.

Optimize the System

Only now are we ready to use the offline evaluation signal produced by the LLM Judge to improve the actual system. We might change the model, improve the retrieval system, or modify the training data.

Each time, we use the same evaluation set and Judge to compare the previous system with the new one.

\[S_0 \rightarrow \text{Change} \rightarrow S_1 \rightarrow \text{Offline Eval}\]

If the Judge score improves, we move to the next iteration.

In this way, the LLM Judge is not merely a tool that evaluates whether a model is good. It becomes a measurement component inside the system-optimization loop.

Close the Loop with Online Evaluation

But the process cannot end there. As discussed earlier, alignment between the LLM Judge and human experts does not guarantee that the signal is aligned with actual user utility.

Once the system is deployed and user signals become available, we need to examine the relationship between offline evaluation and online outcomes.

\[\begin{gathered} \text{Offline Eval} \rightarrow \text{System Optimization} \\ \rightarrow \text{Deployment} \rightarrow \text{Online Evaluation} \end{gathered}\]

Then we ask the most important question:

\[\text{Offline Eval} \uparrow \quad \stackrel{?}{\Longrightarrow} \quad \text{Online Outcome} \uparrow\]
  • Did the model we judged to be better offline actually improve user satisfaction?
  • When the search-relevance Judge improved, did engagement or task success improve as well?
  • When the chatbot-resolution Judge improved, did the real escalation rate or repeat-contact rate improve?

If not, improving the Judge itself may not be the priority. We need to return to the beginning and revisit whether the business objective, proxy signals, and human rubric were defined correctly. Then, if necessary, we start the iteration again.

Iterate Fast

Finally, all of these iterations should happen quickly and frequently. Do not try to create a perfect objective, a perfect rubric, or a perfect Judge from the start. This is such a familiar principle that I will not spend more time explaining it here.

Closing the Loop

Ultimately, when I use LLM-as-a-Judge, the Judge itself is not the most important part.

First, understand the business and the product, and define what we are ultimately trying to optimize.

Next, find measurable signals that serve as useful proxies for that objective. Prefer programmatic or directly observable signals when they are available.

If those signals cannot capture an important quality, and the remaining quality is subjective but trained humans can judge it relatively consistently, then consider an LLM Judge.

Before building the Judge, align the humans. Use human disagreements to discover the decision boundary and improve the rubric. Use the resulting human-labeled dataset to align the Judge with the humans.

Then use the Judge’s offline evaluation signal to optimize the system.

Finally, once the system meets real users, verify that the offline signal is aligned with the online product outcome.

\[\boxed{ \begin{gathered} \text{Business Objective} \rightarrow \text{Measurable Signals} \rightarrow \text{Human Alignment} \\ \rightarrow \text{LLM Judge Alignment} \rightarrow \text{System Optimization} \\ \rightarrow \text{Online Validation} \rightarrow \text{Repeat} \end{gathered} }\]

To me, LLM-as-a-Judge is less a tool that scores outputs instead of a human and more an ML system component that makes part of an otherwise subjective and difficult-to-optimize product objective measurable.

Only when we continue iterating while checking that this measurement remains aligned with the direction we actually want can an LLM Judge play a meaningful role in continuously improving a system.


If you have read this far, something in this post probably caught your interest. If you would like to know more about me, feel free to connect with me on LinkedIn.

References

  1. Part 1: No Data, No Ground Truth: How Hyperconnect Tamed an LLM and Part 2: A Policy-Following Evaluator: LLM-as-a-Judge. Both posts are written in Korean, but I recommend reading them with translation. ↩

  2. Hamel Husain, Using LLM-as-a-Judge for Evaluation: A Complete Guide. His evaluation framework has been a major source of inspiration for me. ↩

  3. Jason Wei, Asymmetry of Verification and Verifier’s Rule. ↩

  4. Hamel Husain, “It’s Hard to Eval” Is a Product Smell. His point is slightly different from the one I emphasize here: if developers find it difficult to evaluate the quality of an output, actual users are also likely to find it difficult to determine whether that output is correct. I borrowed the phrase because I agree with the underlying concern. ↩