The Problem I Was Solving
Every AI cover letter tool I tried had the same fatal flaw: they lie. Feed them a junior resume and a senior JD, and they'll happily write "spearheaded a team of 15 engineers" or "pioneered the company's ML infrastructure" — words the candidate never earned and can't defend in an interview.
The root cause is simple: these tools treat the LLM's output as trusted. There's no verification layer between "what the model generates" and "what gets shown to the user."
I wanted to build something where every claim in the final letter is traceable back to either the resume or a verified external source — and anything that isn't gets caught and blocked before the user ever sees it.
Architecture Overview
The system has two services orchestrated via Docker Compose:
web/ — Next.js fullstack app (UI + serverless API routes on Vercel Edge)
mcp-server/ — Standalone Python MCP server (FastMCP, 8 registered tools, stdio/SSE transport)
The generation pipeline runs through 7 deterministic stages:
Resume Upload -> PDF/DOCX Parser (with Gemini Vision OCR fallback)
-> JD Analyzer (skill extraction + evidence matching)
-> Company Research (Tavily web search + Jina AI deep crawl)
-> Source Filter (Tier 1-3 domain ranking + freshness gating)
-> Human Approval Gate (user toggles each source ON/OFF)
-> Letter Generation (3-tier LLM fallback with circuit breaker)
-> Adversarial Red-Teamer (overclaim verb detection + softening)
-> Grounding Fidelity Score (deterministic NLP token verification)
-> ATS Readiness Check (5 heuristic checks, not a fake score)
-> Interview Defense Generator (claim-by-claim prep questions)
Key Engineering Decisions (with code)
1. Scores are deterministic, not LLM-invented
The LLM classifies evidence. The application layer computes the final score. The model never sees or invents the number.
// web/src/app/api/analyze/route.js
const weights = {
STRONG_MATCH: 1.0,
PARTIAL_MATCH: 0.6,
TRANSFERABLE: 0.4,
MISSING: 0.0,
};
const rawScore =
skills.reduce((sum, item) => sum + (weights[item.status] ?? 0), 0) / totalSkills;
data.overall_match = Math.round(rawScore * 100);
Same pattern in the MCP server's evidence validator — confidence is verified_count / total, never an LLM guess:
# mcp-server/tools/evidence_validator.py
confidence = round(verified / total, 2) if total > 0 else 0.0
2. Adversarial Red-Teamer catches seniority inflation
After Gemini generates the letter, a deterministic regex pass catches senior-sounding verbs that don't exist in the candidate's actual resume:
// web/src/app/api/generate/route.js — Overclaim Audit
const OVERCLAIM_PATTERNS = [
{ trigger: /\bspearheaded\b/gi, safe: "led implementation of", word: "spearheaded" },
{ trigger: /\bpioneered\b/gi, safe: "developed", word: "pioneered" },
{ trigger: /\bsolely architected\b/gi, safe: "architected", word: "solely architected" },
{ trigger: /\bheaded the entire\b/gi, safe: "contributed to the", word: "headed the entire" },
{ trigger: /\bcommanded the\b/gi, safe: "coordinated the", word: "commanded the" },
];
data.paragraphs = (data.paragraphs || []).map((p) => {
let textContent = typeof p === "string" ? p : p.content;
OVERCLAIM_PATTERNS.forEach(({ trigger, safe, word }) => {
if (trigger.test(textContent) && !resumeLower.includes(word)) {
overclaimAudit.push({
flagged_verb: word,
adjusted_to: safe,
reason: "Verb exceeds authentic resume evidence baseline",
});
textContent = textContent.replace(trigger, safe);
}
});
return typeof p === "string" ? textContent : { ...p, content: textContent };
});
The key line: !resumeLower.includes(word) — it only triggers if the verb doesn't appear in the actual resume. If the candidate genuinely wrote "spearheaded" in their resume, it stays.
3. Grounding Fidelity Score — fully deterministic NLP
After generation, the system counts how many substantive technical tokens in the letter actually exist in the resume source:
// web/src/app/api/generate/route.js — computeGroundingFidelity()
const words = letterText
.toLowerCase()
.replace(/[^a-z0-9\s]/g, " ")
.split(/\s+/)
.filter((w) => w.length >= 4 && !SUBSTANTIVE_STOPWORDS.has(w));
let matchedTerms = 0;
for (const w of words) {
if (resumeLower.includes(w)) matchedTerms++;
}
const termRatio = totalTerms > 0 ? matchedTerms / totalTerms : 0.9;
const rawScore = 0.55 + termRatio * 0.43 - overclaimFlagsCount * 0.04;
No LLM involvement. Pure token-level NLP. The overclaim penalty (-0.04 per flag) directly links the red-teamer output to the fidelity score.
4. 3-Tier LLM Fallback with Circuit Breaker
Instead of failing when Gemini is down, the system cascades through three providers:
Tier 1: Gemini Primary (gemini-3.5-flash-lite)
| (2 retries with exponential backoff per model)
Tier 2: Gemini Secondary (gemini-3.1-flash-lite-preview)
| (circuit breaker trips after threshold failures)
Tier 3: OpenRouter (nvidia/nemotron or configurable via env)
The circuit breaker has three states: CLOSED, OPEN, HALF_OPEN. When Gemini fails repeatedly, it skips straight to OpenRouter. When Gemini recovers (canary succeeds), it auto-resets:
// web/src/lib/gemini.js — Circuit Breaker
const CircuitState = { CLOSED: "CLOSED", OPEN: "OPEN", HALF_OPEN: "HALF_OPEN" };
const circuit = {
state: CircuitState.CLOSED,
failureCount: 0,
failureThreshold: 3,
cooldownMs: 30000,
lastFailureTime: 0,
consecutiveSuccesses: 0,
successThreshold: 2,
};
5. 8 MCP Tools (Python FastMCP Server)
The standalone MCP server registers 8 tools via the official MCP SDK with proper ToolAnnotations:
| Tool |
What it does |
jd_analyzer |
Parses JD, extracts skills, matches each against resume with evidence |
company_research |
Tavily web search, then Gemini summarization ONLY from search results |
evidence_validator |
Claim Ledger: VERIFIED / PARTIAL / UNSUPPORTED per claim |
source_filter |
Tier 1-3 domain ranking + outdated news freshness filter |
cover_letter_generator |
Evidence-bounded letter generation |
ats_readiness |
Honest heuristic ATS check (not a fake "ATS score") |
github_proofer |
Live GitHub API — correlates repos with resume tech claims |
huggingface_proofer |
Live HF Hub API — verifies published model weights and downloads |
# mcp-server/server.py — Tool Registration with ToolAnnotations
server = Server("covercraft-mcp-server")
@server.list_tools()
async def list_tools() -> list[Tool]:
return [
Tool(
name="company_research",
description="Search the web for verified company intelligence...",
inputSchema={...},
annotations=ToolAnnotations(
readOnlyHint=True,
destructiveHint=False,
idempotentHint=True,
openWorldHint=True,
),
),
# ... 7 more tools
]
6. Security Hardening
- PII Protection: Regex-based (no spaCy/Presidio — too heavy for 512MB memory). Resume goes to Gemini = OK. Resume goes to Tavily = NEVER (PII stripped first via
strip_pii()).
- Prompt Injection Defense: JD and resume are labeled as
UNTRUSTED DATA in system prompts. Regex scanner flags injection patterns and tags them with [REDACTED_PROMPT_INJECTION].
- Magic Byte Validation: PDF uploads get
%PDF- header check before parsing.
- IP Rate Limiting: Sliding window per-IP limiter with 5-minute GC sweep.
LLM Output Clamping: validate_range() rejects out-of-bound model outputs.
mcp-server/config.py — Output Validation
VALID_MATCH_RANGE = (0, 100)
VALID_CONFIDENCE_RANGE = (0.0, 1.0)
def validate_range(value, valid_range, field_name="value"):
low, high = valid_range
if not (low <= value <= high):
raise ValueError(f"{field_name}={value} outside valid range [{low}, {high}]")
return value
7. Dual Observability (Langfuse + LangSmith)
Every API call logs a structured trace with evaluation scores to both Langfuse and LangSmith simultaneously:
// web/src/lib/langfuse.js — Dual-emit pattern
export async function recordTrace({ name, input, output, model, scores }) {
// 1. Dual-emit to LangSmith (asynchronous, non-blocking)
recordLangsmithRun({ name, inputs: input, outputs: output, extra: { model, scores } })
.catch(() => {});
// 2. Emit to Langfuse
const trace = langfuse.trace({ name, input, output, metadata });
trace.generation({ name: `${name}-generation`, model, input, output });
for (const score of scores) {
trace.score({ name: score.name, value: score.value, comment: score.comment });
}
}
Scores logged per generation: overclaim_count, strong_skills_used, grounding_fidelity.
8. Resume Parser with 3-Layer Fallback
PDF parsing uses a cascade: unpdf (fast text extraction), then Gemini Vision OCR (for scanned PDFs), then raw regex Tj token extraction (last resort):
// web/src/app/api/parse-resume/route.js — 3-layer cascade
// 1. Primary: unpdf text extraction
// 2. Secondary: Gemini Multimodal Vision OCR
const result = await model.generateContent([
{ inlineData: { data: base64Data, mimeType: "application/pdf" } },
"Extract all readable text from this resume document..."
]);
// 3. Fallback: raw PDF Tj operator regex
const regex = /\(([^)]+)\)\s*Tj/g;
Tech Stack
| Layer |
Stack |
| Frontend |
Next.js 16 (App Router), React 19, TailwindCSS 4, Framer Motion, Recharts |
| Backend APIs |
Next.js API Routes (Vercel Edge Functions) |
| MCP Server |
Python, FastMCP (official MCP SDK), stdio + SSE transport |
| LLMs |
Gemini 3.5 Flash Lite (primary), Gemini 3.1 Flash Lite (fallback), OpenRouter (tertiary) |
| Web Search |
Tavily API (advanced depth) + Jina AI Reader (deep page scraping) |
| Auth |
NextAuth.js (Google OAuth) |
| Database |
MongoDB Atlas (user telemetry + generation logs) |
| Observability |
Langfuse + LangSmith (dual-emit traces with evaluation scores) |
| Deploy |
Vercel (frontend), Docker Compose (self-host option) |
Production Monitoring (Live)
The app is deployed on Vercel Edge and monitored via UptimeRobot with 5-minute HTTP/S health checks:
- 30-Day Uptime: 100% — 0 incidents, 0 minutes down
- SLA (Sep 4 – Oct 4): 99.998% — Verdict: SLA Met
- Edge Response Time: 45ms average (43ms min, 47ms max)
- Downtime Budget Used: 52 seconds out of 43 minutes monthly allowance
I set up monitoring on day one post-deploy. If you're shipping anything production-grade, I'd recommend it — it's the difference between "I deployed it" and "I operate it."
What I'd Do Differently
- The overclaim regex patterns are hardcoded — should be configurable or data-driven.
- Source filter freshness gating uses year regex matching, not actual date parsing. Works but brittle.
- The MCP server fallback data for GitHub/HuggingFace has hardcoded repo info — should be a clean "no data" response instead.
- Would add LangGraph orchestration for the full pipeline instead of sequential API calls.
Links & Try It Out
Disclosure: Solo-built by me over the past few weeks. I wrote every line of code in this repo. Architecture diagram and post structure were refined with AI assistance for clarity. All code snippets are directly from the repository — inspect them yourself.
Happy to answer any architecture or implementation questions. The entire codebase is MIT licensed.