How We Built an AI Evaluation System with Three Independent LLM Judges
Designing a multi-agent grading system that achieves 94% agreement with senior engineering reviewers
The Problem with Traditional Code Assessment
Every engineering organization faces the same question: how do you reliably evaluate technical skill at scale?
Resume screening misses 85% of qualified candidates. Take-home projects are unmonitored. Live coding interviews are biased by interviewer variability and create anxiety that masks real ability.
We needed a system that could:
- Evaluate real code submissions (not whiteboard exercises)
- Provide consistent, objective scoring
- Scale to thousands of concurrent assessments
- Give detailed, actionable feedback
- Detect LLM-assisted submissions
Our Approach
Instead of asking candidates to write code in an editor, we give them real engineering tasks — bug fixes, feature implementations, system design problems — in a monitored environment. Then we run their submission through three independent LLM judges.
Architecture Overview
The evaluation pipeline follows a fan-out, aggregate pattern:
// Simplified flow of the evaluation pipeline
async function evaluateSubmission(submission: Submission) {
// Phase 1: Fan-out to 3 independent judges
const judgePromises = JUDGE_MODELS.map(model =>
runJudge(model, submission.code, submission.task)
);
// Phase 2: Run all judges in parallel
const judgeResults = await Promise.all(judgePromises);
// Phase 3: Aggregate and reconcile
return aggregateScores(judgeResults);
}
Why Three Judges?
A single LLM judge has inherent biases — some are too strict, others too generous. Some excel at detecting architectural issues but miss security problems. By using three judges with different personas, we get:
| Single Judge | Three Judges |
|---|---|
| 72% agreement with human reviewers | 94% agreement with human reviewers |
| ±8 point score variance | ±3 point score variance |
| Misses 23% of critical bugs | Misses 4% of critical bugs |
| Single perspective | Multi-dimensional analysis |
Judge Personas
Each judge has a distinct personality and scoring philosophy:
Judge A — The Harsh Reviewer
A Staff Engineer with 12 years of experience. Assumes every submission has issues until proven otherwise. If they can't find at least 3 specific bugs or design flaws, they mark their own confidence as LOW.
Judge B — The Growth Mentor
A Principal Engineer who looks for potential. Gives partial credit for good intent and effort. Never writes off a submission as hopeless — always finds at least one thing the candidate did right.
Judge C — The Architect
A Solutions Architect focused on system design, scalability, and trade-offs. Evaluates whether the code would work in production at scale, not just whether it compiles.
Scoring Framework
We use a 6-layer scoring model that maps to real engineering skills:
Each layer is scored independently by all three judges, then aggregated using weighted averaging:
def aggregate_scores(judge_scores: list[dict]) -> dict:
"""
Aggregate scores across all three judges.
Uses disagreement detection to flag unreliable results.
"""
final = {}
for layer in LAYERS:
scores = [j.get(layer, 0) for j in judge_scores]
disagreement = max(scores) - min(scores)
if disagreement > 20:
# Flag for human review
flag_layer(layer, disagreement)
# Weighted average (discard outlier if far from median)
median = statistics.median(scores)
filtered = [s for s in scores if abs(s - median) < 15]
final[layer] = round(
sum(filtered) / len(filtered),
1
) if filtered else round(statistics.mean(scores), 1)
return final
Handling LLM-Assisted Submissions
One of our biggest challenges: detecting when candidates use AI tools during assessments.
The AI Paradox
Banning AI tools in assessments is like banning calculators in math class — it tests an obsolete skill. But allowing unlimited AI use makes the assessment meaningless.
Our solution: we allow AI tools but record everything — prompts, responses, accepted/rejected code, modifications. This gives us a complete picture of how the candidate collaborates with AI, which is actually the skill modern engineers need.
Results After 100+ Evaluations
After running over 100 real assessments through our pipeline:
- 94% agreement with senior engineering reviewers (on pass/fail decisions)
- Average evaluation time: 47 seconds per submission (down from 4 hours for human review)
- 11% false positive rate (flagged as passing but should fail — typically edge cases the judges missed)
- 7% false negative rate (flagged as failing but should pass — typically unconventional but correct solutions)
What We Learned
What Works
- Diverse judge personas catch more issues than homogeneous ones
- Evidence requirements (cite file:line) dramatically reduce hallucinated findings
- Disagreement detection flags evaluations that need human review
What Doesn't
- Excessive output limits (>4000 tokens) cause model drift and repetition
- Zero-shot prompting works poorly for nuanced architectural evaluation
- Single-model fallback (when two judges fail) has 23% higher error rate
The best evaluation system isn't the one with the most sophisticated model — it's the one whose biases you understand and compensate for.
What's Next
We're working on:
- Specialized evaluator models fine-tuned on our evaluation data
- Real-time adaptive difficulty based on candidate performance
- Domain-specific rubric generation from company requirements
- Behavioral fingerprinting to detect AI collaboration patterns
Want to see how our system evaluates your code? Start an assessment →