Skip to content
All work

Case study 03 / 11

Enterprise workforce and HR

DeliveredClient work · NDA

Adaptive AI-Readiness Assessment Platform

As the sole engineer, I designed and built an adaptive platform that measures the AI skills of a large enterprise's workforce. Questions adjust to each answer, AI grades written responses under strict safeguards, and any uncertain grade goes to a human reviewer. It meets WCAG 2.2 AA accessibility standards and passed its load test without a single failed request.

System

Tested rules decide

  • Which question comes next
  • Multiple-choice grading
  • Course selection and learning path
AI stays out of the decisions

AI

AI helps, within limits

  • Phrasing questions
  • Grading written answers against a rubric
  • Writing the report text
  • Flagging possible AI-written answers
Any grade the system isn't confident about goes to a human reviewer.

The challenge

The client measured its employees' AI skills with a static online form: fixed multiple-choice questions, points-based scoring and results in a spreadsheet. It couldn't tell a beginner from an expert, couldn't judge reasoning, and gave nobody a clear next step.

They wanted a system that:

  • assesses every employee across a framework of AI skills and job roles
  • adapts the questions to each person in real time
  • grades written answers with AI without handing decisions to it
  • recommends courses from their own catalogue
  • gives HR and learning teams review tools, reporting and a full audit trail
  • scales to thousands of employees and meets accessibility and privacy standards

What I delivered

  • The employee experience. A simple invite-only sign-in and a clear consent notice, then an assessment that gets harder after a right answer and easier after a wrong one, with a written question at the top to identify the strongest people. Employees can answer by voice. At the end they get an instant report and a personal learning path that puts the most important gaps first and skips courses they've already completed.
  • The reviewer workspace. A queue of AI grades the system wasn't confident about, plus possible integrity issues. When a reviewer changes a grade, the result is recalculated so everything stays consistent.
  • The admin console. User invitations, the skills framework, role targets, the course catalogue, assessment campaigns with automatic reminders, team reporting that never shows groups smaller than five people, and an audit log.
  • Integrity signals. Answers that come too fast, copy and paste, switching away from the page, and possible AI-written answers are flagged for a person to look at. Nothing is decided automatically.

How it works

The most important decision was keeping AI out of the decisions. Clear, tested rules handle everything that matters: which question comes next, multiple-choice grading, course selection and the learning path. AI does four narrow jobs:

  1. Phrasing questions naturally.
  2. Grading written answers against a rubric. This is the only AI task that sees the rubric, and its output never goes to the employee.
  3. Writing the report text, using only facts the system has already confirmed.
  4. Flagging answers that may have been written by AI.

Key decisions

  • A safety net for every AI step. Every AI response is checked. If the AI service is unavailable, grades go to human review and reports are produced from templates, so the platform keeps working.
  • Majority vote for borderline grades. AI grading wasn't perfectly consistent, so borderline answers are graded three times and decided by majority. Any disagreement goes to a person.
  • Protection against manipulation. Employee answers are sealed off so they can't trick the AI grader, and suspicious attempts are sent to a human.
  • No invented content. The report is checked so it can't mention made-up courses, links or numbers, or anything that sounds like a hiring decision.
  • Kept deliberately simple. I didn't add extra AI frameworks the problem didn't need, which kept the system predictable and easy to audit.

Built for trust

  • Four roles with access checked on every single action. An automated test fails the build if any feature is left without an access rule.
  • Secure passwords, sessions that can be revoked, and protection against repeated login attempts.
  • A permanent audit log. If an action can't be recorded, it doesn't happen.
  • Privacy by design: raw written answers and voice recordings are never stored, and a test guarantees email addresses never reach the AI.
  • Ready for company-wide single sign-on.

The results

The platform replaces a static form with an adaptive assessment that tells a beginner from an expert on each skill, grades reasoning as well as knowledge, and gives every employee a personal learning path, with every AI judgement reviewable by a person and every action on record.

Under the hoodShow technical details
  • Engineering judgement: in week one I compared the build against the requirements and found six missing items, then built most of them. I declined to build AI-based recruitment screening without consent, fairness testing and an impact assessment, and wrote a responsible-AI proposal instead. I also kept all 22 client decisions in a decision log.
  • AI: Azure OpenAI with structured outputs and code validation, versioned prompts, and a calibration tool that compares AI grades with human grades.
  • Scaling: it runs in containers on Azure. The service keeps no session state in memory, so more servers can be added freely. Containers run with minimal permissions, and every release is checked automatically for leftover secrets or test code.
  • Performance: load testing with k6 showed the limit was processing power, not the database, so I kept the database settings as they were. I also found and fixed issues in the load test itself.
  • Quality: 830 automated tests covering every path through the assessment, and a full assessment completed using only a keyboard.