Back to blog
Engineering2026-06-155 min read

How We Built an AI Evaluation System with Three Independent LLM Judges

Designing a multi-agent grading system that achieves 94% agreement with senior engineering reviewers

A
Al pacino
Agentic AI Engineer, Taqeem
#AI#LLM#System Design#Evaluation#Architecture

The Problem with Traditional Code Assessment

Every engineering organization faces the same question: how do you reliably evaluate technical skill at scale?

Resume screening misses 85% of qualified candidates. Take-home projects are unmonitored. Live coding interviews are biased by interviewer variability and create anxiety that masks real ability.

We needed a system that could:

  • Evaluate real code submissions (not whiteboard exercises)
  • Provide consistent, objective scoring
  • Scale to thousands of concurrent assessments
  • Give detailed, actionable feedback
  • Detect LLM-assisted submissions
→

Our Approach

Instead of asking candidates to write code in an editor, we give them real engineering tasks — bug fixes, feature implementations, system design problems — in a monitored environment. Then we run their submission through three independent LLM judges.

Architecture Overview

The evaluation pipeline follows a fan-out, aggregate pattern:

// Simplified flow of the evaluation pipeline
async function evaluateSubmission(submission: Submission) {
  // Phase 1: Fan-out to 3 independent judges
  const judgePromises = JUDGE_MODELS.map(model =>
    runJudge(model, submission.code, submission.task)
  );

  // Phase 2: Run all judges in parallel
  const judgeResults = await Promise.all(judgePromises);

  // Phase 3: Aggregate and reconcile
  return aggregateScores(judgeResults);
}

Why Three Judges?

A single LLM judge has inherent biases — some are too strict, others too generous. Some excel at detecting architectural issues but miss security problems. By using three judges with different personas, we get:

Single Judge Three Judges
72% agreement with human reviewers 94% agreement with human reviewers
±8 point score variance ±3 point score variance
Misses 23% of critical bugs Misses 4% of critical bugs
Single perspective Multi-dimensional analysis

Judge Personas

Each judge has a distinct personality and scoring philosophy:

ℹ

Judge A — The Harsh Reviewer

A Staff Engineer with 12 years of experience. Assumes every submission has issues until proven otherwise. If they can't find at least 3 specific bugs or design flaws, they mark their own confidence as LOW.

⚠

Judge B — The Growth Mentor

A Principal Engineer who looks for potential. Gives partial credit for good intent and effort. Never writes off a submission as hopeless — always finds at least one thing the candidate did right.

✓

Judge C — The Architect

A Solutions Architect focused on system design, scalability, and trade-offs. Evaluates whether the code would work in production at scale, not just whether it compiles.

Scoring Framework

We use a 6-layer scoring model that maps to real engineering skills:

0-100
Code Quality
0-100
Architecture
0-100
Reasoning
0-100
Security & Perf
0-100
Communication
0-100
Seniority

Each layer is scored independently by all three judges, then aggregated using weighted averaging:

def aggregate_scores(judge_scores: list[dict]) -> dict:
    """
    Aggregate scores across all three judges.
    Uses disagreement detection to flag unreliable results.
    """
    final = {}
    for layer in LAYERS:
        scores = [j.get(layer, 0) for j in judge_scores]
        disagreement = max(scores) - min(scores)

        if disagreement > 20:
            # Flag for human review
            flag_layer(layer, disagreement)

        # Weighted average (discard outlier if far from median)
        median = statistics.median(scores)
        filtered = [s for s in scores if abs(s - median) < 15]
        final[layer] = round(
            sum(filtered) / len(filtered),
            1
        ) if filtered else round(statistics.mean(scores), 1)

    return final

Handling LLM-Assisted Submissions

One of our biggest challenges: detecting when candidates use AI tools during assessments.

⚠

The AI Paradox

Banning AI tools in assessments is like banning calculators in math class — it tests an obsolete skill. But allowing unlimited AI use makes the assessment meaningless.

Our solution: we allow AI tools but record everything — prompts, responses, accepted/rejected code, modifications. This gives us a complete picture of how the candidate collaborates with AI, which is actually the skill modern engineers need.

Results After 100+ Evaluations

After running over 100 real assessments through our pipeline:

  • 94% agreement with senior engineering reviewers (on pass/fail decisions)
  • Average evaluation time: 47 seconds per submission (down from 4 hours for human review)
  • 11% false positive rate (flagged as passing but should fail — typically edge cases the judges missed)
  • 7% false negative rate (flagged as failing but should pass — typically unconventional but correct solutions)
94%
Agreement with Humans
47s
Average Eval Time
11%
False Positive Rate
7%
False Negative Rate

What We Learned

What Works

  1. Diverse judge personas catch more issues than homogeneous ones
  2. Evidence requirements (cite file:line) dramatically reduce hallucinated findings
  3. Disagreement detection flags evaluations that need human review

What Doesn't

  1. Excessive output limits (>4000 tokens) cause model drift and repetition
  2. Zero-shot prompting works poorly for nuanced architectural evaluation
  3. Single-model fallback (when two judges fail) has 23% higher error rate

The best evaluation system isn't the one with the most sophisticated model — it's the one whose biases you understand and compensate for.

What's Next

We're working on:

  • Specialized evaluator models fine-tuned on our evaluation data
  • Real-time adaptive difficulty based on candidate performance
  • Domain-specific rubric generation from company requirements
  • Behavioral fingerprinting to detect AI collaboration patterns

Want to see how our system evaluates your code? Start an assessment →

For Developers

Prove your engineering skills

Take a real coding evaluation, skip resume screens, and get discovered by top hiring teams.

Get Verified for Free

100% Free · Privacy Protected

For Hiring Managers

Hire pre-vetted senior engineers

Access vetted candidate profiles with multi-judge AI scoring and verified code samples.

Browse Verified Talent

Zero Placement Fees · Instant Access