CustomLabs
Evaluation

LLM-as-Judge

LLM-as-judge is an evaluation technique that uses a — typically stronger or differently-configured — LLM to score another model's outputs against a rubric, at a scale human review can't match. It's useful for grading subjective qualities like tone, relevance, or faithfulness to a source document, but it inherits the judging model's own biases and blind spots, so it's normally calibrated against a smaller human-labeled sample rather than trusted blind. Treat it as one signal in an eval suite, not the whole suite.

← Back to the full glossary

navigate select esc close