1 / 8
LLM Eval Agent architecture

LLM Safety as a Quality Gate

A GitHub App that enforces LLM safety standards at the PR level
the same way linters enforce code quality.

While LLM evaluation frameworks (EleutherAI Harness, DeepEval, HELM) and ML fairness libraries (AIF360, RAI Toolbox) exist independently, no prior work integrates automated bias/robustness evaluation as a native GitHub App with PR-level merge enforcement, confidence scoring, and longitudinal audit logging.

Kavita Jadhav

K11 Software Solutions LLC · Texas, United States

⬥ App Repo ⬥ Demo Repo
The Problem

LLM failures go undetected until production

Bias

Demographic bias in training data — invisible to functional tests

🎯
Fairness

Imbalanced accuracy across groups — missed at code review

🛡
Robustness

Brittleness under adversarial inputs — surfaces only at scale

⚠ No existing tool integrates LLM safety into PR review as a blocking merge gate

Our Solution

Shift-left LLM safety into every pull request

🔔
PR Opened
GitHub webhook
🔒
Auth + Trigger
HMAC · JWT
🧪
LangTest Eval
Async · 21 sec
Check Run
Block or merge
📊
Audit Log
Longitudinal

Like linters enforce code quality — LLM Eval Agent enforces safety standards at merge time

Novel Contributions

6 engineering innovations

🔗
GitHub App + PR merge gate
Block merge when safety thresholds fail
🤖
Adversarial red-teaming in CI/CD
Typo injection · spelling-variant perturbation
📊
Confidence score reporting
Avg / Min / Max surfaced in Check Run
📋
Structured audit log + trend API
Per-commit JSONL · /trend REST endpoint
🕐
Scheduled drift detection
APScheduler cron · no PR required
🧾
109-test validation suite
All platform components independently tested
Evaluation Results

3 architectures · 5 test categories

DistilBERT BERT-base RoBERTa
0% 25% 50% 75% 100% 98% 100% 98% 98% 100% 98% 86% 94% 97% 100% 82% 100% FAIL
Test Category
① Bias: female sentiment
98–100% PASS
② Bias: male sentiment
98–100% PASS
③ Fairness: gender F1
0% FAIL
④ Robustness: typos
86–97% PASS
⑤ Robustness: spelling
82–100% PASS

Gate thresholds: Bias ≥ 80% · Fairness ≥ 80% · Robustness ≥ 75%

Headline Finding

Fairness: 0% across ALL architectures

0%
DistilBERT
FAIL
=
0%
BERT-base
FAIL
=
0%
RoBERTa
FAIL

This is not an architecture defect. All three models were fine-tuned on SST-2, which contains gender-correlated sentiment patterns. Remediation requires dataset-level intervention — not model swapping.

✓ Surfaced automatically on the first PR — exactly the shift-left outcome intended

Live Demo · PR #4 · kavitaj11/llm-eval-demo

End-to-end in 21 seconds

0%
Bias gate
PASS
threshold 80%
0%
Robustness gate
PASS
threshold 75%
0%
Avg confidence
INFO
min 55% · max 100%
21s evaluation time
🔒Block-merge on timeout ✓
🚀Railway · auto HTTPS
55% min confidence flagged
Open Source · Apache 2.0

Available now as an installable GitHub App

Installs on any repository. No infrastructure to manage. Safety evaluation automatically on every PR.

⬥ App Source Code ⬥ Live Demo Repo
27
References
109
Unit tests
6
Novel contributions
🔗 DOI: 10.5281/zenodo.20823852 — © 2026 Kavita Jadhav · Apache 2.0

Also available on Zenodo · ResearchGate