HR & People automation · Multi-agent decision support
Hiring Panel
Six AI agents review one synthetic candidate for one role and debate it live. Every claim must quote the candidate’s documents, a fairness auditor challenges the others, and code computes the scores and the offer range. The panel never decides: it prepares a brief for a person.
1 · Choose a role and a candidate
2 · Panel debate
- Press “Convene the panel” to watch six agents debate. Each turn is one request; the server checks every quote and statement before the scores move.
3 · Brief for the hiring manager
Profile against the role
Offer range (synthetic benchmark)
Points to verify at interview
Dissent log
Audit trail
Human decision
The panel’s outcome is advice. You decide. Your choice is shown on this page only: nothing is stored or sent.
How it works
- Orchestrated by code: the page drives a fixed turn order (three opening assessments, an audit with responses, then pay band and brief) with a hard cap of nine turns. Each turn is one request to a Cloudflare Pages Function.
- One agent, one compact prompt per turn: the Function builds a short prompt for DeepSeek (OpenAI-compatible API, small output limit, thinking disabled) from allow-listed data and the run state. Visitors only pick a role and a candidate, so no free text reaches the model.
- No server session: the run state travels with each request, signed with an HMAC. The server validates its structure and signature before acting, so altered state is refused.
- Guardrails in code, not in the prompt: every quote is matched against the cited document (case and whitespace normalised), a list of protected characteristics and proxies strikes statements from the record, and pay figures that the policy did not compute are struck.
- Never fails because of the LLM: if the key is missing, the provider errors or a turn is invalid twice, the page replays a recorded live run through the same checks and scoring, and the badge says “Recorded run”.
Scoring formula (agents propose, code aggregates)
- mean(r, a) = mean of agent a's accepted scores (1 to 5) for requirement r, or 0 if none
- Role fit = 20 × Σ w_r · mean(r, Hiring Manager) ÷ Σ w_r
- Technical depth = 20 × Σ w_r · mean(r, Technical Assessor) ÷ Σ w_r, over technical requirements
- Growth potential = 20 × mean of the People Partner's accepted scores
- Evidence strength = 100 × (½ · accepted ÷ claims made + ½ · weighted share of requirements with an accepted score of 3 or more)
- Ramp-up weeks = 2 + round(10 × (1 − Σ w_r · best_r ÷ 5 ÷ Σ w_r)); ramp-up score = 100 × (12 − weeks) ÷ 10
- Offer: floor = P25; ceiling = min(P75, equity cap); recommended = floor + (ceiling − floor) × clamp((role fit − 50) ÷ 50, 0, 1), rounded to £500. Previous pay is never an input.
Responsible use
Using AI to evaluate job candidates is a high-risk use under the EU AI Act (Annex III, employment), so this demo is built around the controls such a system needs. Human oversight: the panel can only suggest “Advance to interview” or “Needs more evidence”, there is no automated rejection, and a person records the decision. Evidence traceability: every scored claim carries a quote that code verifies against the cited document, and unverified claims are excluded. Bias checks: the Fairness and Compliance Auditor challenges the panel, and a code-side filter strikes any statement that relies on age, gender, family status, career breaks, nationality, appearance, a name or previous pay. Logging: every claim, revision, strike and token count is kept in the run record and shown in the audit trail. All candidates, the company and the pay benchmark are synthetic.
Source code and tests: github.com/nepryoon/hr-hiring-panel (opens in a new tab)