Definition
Using a language model to evaluate, compare, classify, or score outputs produced by models or agents.
Distinguish it from nearby terms
LLM judging is scalable approximate evaluation, not independent ground truth.
Check your understanding
Separate producer and judge context, allow abstention, test framing sensitivity, and combine with deterministic evidence.