Don't hunt through the portfolio. Examine it.

A portfolio you can examine — inspect the work, its evidence, and the limits of what it proves.

I designed and built the product end to end — interaction model, knowledge-base retrieval, typed answers, Jobfitter rubric, review workflow, React/Express implementation.

Key results

A portfolio that can be cross-examined.

The bet

Hiring teams arrive with questions, not time to read every page. Ask directly, inspect the source, or open the full story.

Grounded

Answers use a generated knowledge base of portfolio evidence.

Inspectable

Citations expose the source; missing evidence stays missing.

Allowed to disagree

Jobfitter separates Skills from Role Shape and can return a weak fit.

The obvious build → the bet

“Add a chatbot to the portfolio.”

“Build a new interaction mode with my work.”

The agent answers about the exact line you select on the page.

Ask in context

You don’t have to retype what you’re looking at. Select any line on the page and a small toolbar comes to you — copy it, or ask about it and the agent answers with that exact excerpt as context. No right-click, no menu hunting: the selection is the question.

Answers are interfaces.

How it works

The question determines the response: prose, comparison, chart, source artifact, calculator, or structured Jobfitter card.

Designed, built, and calibrated as a product.

How it's built

I designed the interaction, retrieval, answer contracts, Jobfitter rubric, and review workflow, then shipped the React/Express implementation.

Citations expose the source

A confident answer is only useful if you can check it.

Citations are prompt-required, not structurally validated per string; the interface resolves them to portfolio sources.

Two scores, read separately

One “fit %” would hide whether skills and role shape each match.

The model authors each materiality-weighted score; the server clamps each score, derives the verdict band, and falls back deterministically if one is missing.

Built in the open

“I build AI products” is a claim recruiters discount.

~1 year of daily agentic coding, test-gated changes, and a decision record linking calibration rules to observed failures.

Live; hiring outcomes remain unmeasured.

Where it stands

Real submissions supply calibration cases and failure modes — not yet proof of better hiring decisions or faster screening.

No ground-truth eval set yet

Calibration is hand-checked against real submissions, not benchmarked on a labelled set — that’s next.

Keyword retrieval, by choice

Deterministic, fast, and cheap at this size; vectors are the upgrade when paraphrase misses justify it.

What worked · what’s next

What worked

Separate dimensions make the assessment easier to challenge.

Requirement-level verdicts show how the assessment formed, instead of one flattering score.

What’s next

Build the labelled eval set before expanding the surface.

A labelled set of strong, partial, and weak roles, run on every prompt change.

FAQ

Isn't this just a thin wrapper around an LLM?

The product work sits around the model. It authors materiality-weighted scores and requirement-level verdicts. The server clamps each score, derives the verdict band, and falls back to a deterministic bullet mix when needed. Prompt rules require citations, distinguish missing evidence from contradiction, and block flattering language; typed UI contracts turn the output into an inspectable scorecard. Per-string citation validation is not implemented yet.

Why two scores instead of one fit number?

A single "match %" conflates two different questions. Skills Match asks "can he do this work?"; Role-Shape Match asks "is this the work he's choosing to do next?" Someone can be entirely capable of a role that points away from where they're heading. One number hides that; two numbers surface it. That separation is the core design insight.

Does it ever actually say no?

Yes. In the first five weeks of real use (June–July 2026), 7 of 35 pasted roles came back weak on a dimension and 6 more were refused as not a job description. Real gaps are named, not softened. A scorer that can't decline isn't measuring anything.

Canonical page