LLM-as-judge
Using a second model call to grade the output of the first against a written rubric.
This is how teams evaluate subjective output - tone, helpfulness, faithfulness - at a scale where human grading is impossible.
Always calibrate the judge: grade ten cases yourself and check the judge agrees before trusting its scores.