Skip to content
All work

Case study 04 / 11

HR technology

DeliveredClient work · NDA

AI Talent Assessment Platform

I designed and built, on my own, a hiring platform where a recruiter creates a fair, defensible assessment just by describing the role. An AI assistant drafts the test, candidates take it in a secure environment, and every score comes from established scoring science, never from the AI's opinion. I delivered it in about three months.

AI

AI writes

  • A skills plan from the job description
  • Draft questions and scenarios
  • Suggested answer keys for a person to confirm
AI firewall

System

Proven methods score

  • Every score and level
  • Pass, fail or fit
  • Confirmed answer keys
The same answers always produce the same score.

The challenge

Writing a good hiring test is slow and needs specialist expertise that most HR teams don't have. So companies fall back on generic tests that don't fit the role, or on unstructured interviews that don't hold up to scrutiny.

AI looks like the obvious answer until you try it. It will happily write twenty questions, and just as happily tell you a candidate is "a strong fit". That isn't a decision anyone can defend to a rejected applicant, an auditor or a regulator. The brief was to get the speed of AI without letting it anywhere near the score.

What I delivered

  • An AI assessment assistant. A recruiter describes the role, pastes a job description, or uploads a document or even a photo of a paper test. The assistant suggests the right type and length of assessment, drafts the questions and shows them live as editable drafts. The recruiter refines them just by chatting: "make section two harder", "swap that question".
  • Reliable scoring. I integrated a proven scoring engine and built all the scoring logic around it. It covers ability tests that adapt to each candidate, personality questions designed to resist faking, and workplace judgement scenarios. Results are presented as clear levels.
  • Skills-based matching. A shared vocabulary of thousands of skills and occupations connects job profiles to candidates. Skills that the AI picks out of a job description only count once a recruiter confirms them, and candidates see exactly which skills they have, lack, or need to strengthen.
  • Language testing. Reading, listening, grammar, writing and speaking, aligned to the international CEFR levels. Speaking is scored automatically, recruiters can upload their own tests, and Arabic is supported.
  • The full hiring workflow. Employer requests, a review workspace for staff, secure test-taking, results that stay hidden until deliberately released, PDF reports and email notifications.

How it works: an AI firewall

AI writes. Proven methods score.

AI canAI can never
Turn a job description into a skills planCalculate a score or level
Draft questions and scenariosDecide pass, fail or fit
Suggest an answer key for a person to confirmChange a confirmed answer key
Rate a written answer against a rubricDecide how much that rating counts

The same answers always produce the same score, and every step can be checked.

Key decisions

  • Secure test-taking. Switching tabs or leaving full screen ends the attempt, with a short grace period at the start so candidates aren't failed by the act of entering full screen.
  • Structure over instructions. AI doesn't always follow instructions, so the rules that must hold are enforced by the system, not by the wording of a prompt.
  • Real-world testing. A structured review of real user journeys found 34 issues that automated tests had missed, mostly screens people couldn't reach. Walking through real journeys became a required step before anything is called done.
  • One consistent design system across all three portals, with light and dark themes kept in perfect step.

Built for trust

  • Answer keys, scoring settings and marking rubrics stay on the server and never reach the candidate.
  • Personal details are removed before any text goes to AI, uploaded content is checked for manipulation attempts, and AI-generated content is screened before use.
  • Every organisation's data is kept separate, and every action is logged without storing personal data or AI prompts.
  • The platform can be fully demonstrated and tested without spending anything on AI, using a built-in demo mode.

The results

A complete platform, built and delivered by one engineer. The client can create an assessment by describing it, send it to candidates, have them take it securely, and receive a scored, exportable report where every number comes from methods that are consistent, explainable and fair.

The principle I'd defend anywhere: use AI for the part that is hard for people and easy to check, writing content, and never for the part that has to be consistent, explainable and fair.

Under the hoodShow technical details
  • Scoring science: item response theory with adaptive testing, forced-choice personality models and situational judgement scoring, in a standalone Python package (numpy and scipy) tested against data with known answers.
  • AI: OpenAI behind one interface with real, test and demo versions; a tool-using agent that streams its work live to the screen; image understanding for uploaded tests. I fixed a problem that only appeared with the live AI service by redesigning every response format and adding checks for it.
  • Skills data: the public ESCO skills and occupations taxonomy.
  • Cloud: runs in Docker containers on Azure.
  • How I worked: written design and plan before each feature, tests as a deliverable, parallel workstreams, and documentation (architecture notes, API contract, runbooks and a user manual) so the client's team can run and extend it.