Build a structured interview kit instead of improvising questions

Generate a reusable question set and scoring rubric so every candidate gets the same fair shot.

For anyone who interviews a few times a year · 8 steps · 7 min

How it works today.

You open the CV two minutes before the Zoom link goes live, scan for anything that might spark a question, and wing it. One candidate gets grilled on system design because their last job mentioned microservices. The next gets softballs about team culture because you were running late and needed something easy.

When you sit down to compare notes with the hiring manager, you realise you asked completely different questions. You fall back on gut feel, and the person who happened to mention a technology you like gets the edge. Nobody can defend the decision in writing, and you have no idea whether you are selecting for skill or for small talk.

Before you start.

All of it has to be true, or step one fails in a way that is annoying to debug.

  • A job description or brief that lists the actual responsibilities and must-have skills.
  • Access to a generalist agent (ChatGPT, Claude, Gemini) or a builder tool where you can save and version a prompt.
  • Agreement with anyone else interviewing that you will all use the same question set and rubric.
  • A document or folder where you can store the kit and scoring notes for each candidate.

The steps.

  1. Feed the job description to the agent and ask for interview dimensions

    Paste your job description or a bullet list of the role's core responsibilities. Ask the agent to propose four to six dimensions you should evaluate—things like technical judgment, communication under ambiguity, or collaboration with non-technical stakeholders. Review the list and remove anything that feels like a nice-to-have rather than a dealbreaker.

    Paste this
    Here is a job description:
    
    [paste your JD]
    
    Propose 4–6 concrete dimensions I should evaluate in a structured interview. For each dimension, write one sentence explaining why it matters for this role. Do not include generic traits like 'culture fit'.
  2. Generate two behavioural questions per dimension

    For each dimension the agent proposed, ask it to write two behavioural questions that probe for evidence. Behavioural questions start with 'Tell me about a time…' or 'Describe a situation where…'. Read every question aloud to yourself. If it sounds like a riddle or could be answered with a single sentence, rewrite it or ask the agent to try again.

    Paste this
    For each dimension below, write two behavioural interview questions. Each question should ask the candidate to describe a specific past situation, what they did, and what happened. Avoid hypotheticals.
    
    Dimensions:
    [paste the list from step one]
  3. Build a three-point rubric for each dimension

    Ask the agent to define what a weak, acceptable, and strong answer looks like for each dimension. The rubric should describe observable evidence, not adjectives like 'passionate' or 'proactive'. You will use this rubric during the interview to score answers in real time, so it needs to fit on one page.

    Paste this
    For each dimension, write a three-level rubric: weak, acceptable, strong. Describe what evidence you would hear in the candidate's answer at each level. Be specific enough that two interviewers would score the same answer the same way.
    
    Dimensions:
    [paste the list from step one]
  4. Assemble the kit in a reusable document

    Copy the questions and rubric into a single document with space to write notes under each question. Include a header with the candidate's name, the date, and the role. Save this as a template. Every time you interview someone for this role, you will duplicate the template and fill in their answers. Do not add extra questions mid-interview unless every candidate gets them.

  5. Run a dry interview with a colleague

    Book fifteen minutes with someone on your team who is not involved in hiring. Give them one of the behavioural questions and ask them to answer it as if they were a candidate. Score their answer using the rubric. If you cannot decide between two levels, the rubric is not specific enough—revise it before you interview real candidates.

  6. Conduct the interview and score in real time

    Open the duplicated template during the call. Ask each question in order, take notes on what the candidate says, and assign a score immediately after they finish answering. Do not wait until the end of the interview to score, because you will forget the details. If a candidate gives a thin answer, ask one follow-up to probe for specifics, then move on.

  7. Compare candidates using only the scored rubric

    After you have interviewed everyone, open all the completed templates side by side. Compare scores dimension by dimension. If two candidates have similar scores but you still have a preference, write down the evidence that tips the balance—something you observed and wrote in your notes, not a feeling. If you cannot point to evidence, the preference is bias.

  8. Ask the agent to summarise patterns across all candidates

    Once you have scored three or more candidates, paste the anonymised scores and a few bullet points of notes into the agent. Ask it to identify which dimensions most candidates struggled with and whether your rubric is distinguishing between people. If everyone scores the same on a dimension, that question is not useful.

    Paste this
    Here are anonymised scores and notes from [number] candidates for the same role:
    
    [paste scores and key notes]
    
    Which dimensions separated candidates most clearly? Which dimensions did most candidates score similarly on? Suggest one revision to make the rubric more discriminating.

What you keep.

Automating the typing does not move the accountability. These stay with a person.

  • Deciding which dimensions actually matter for the role, because the agent does not know your team's weak spots or the problems this hire will face in month two.
  • Scoring each answer during the interview, because the agent was not in the room and cannot hear tone, recovery from a fumble, or whether the candidate is reciting a script.
  • Making the final hiring decision by weighing the evidence you collected, not by outsourcing judgment to a tool that has never worked with any of these people.

Once it works.

The first run is the demo. These are where the time actually comes back.

Run it on a loop

After every hiring round, feed the agent your notes on which questions produced the most signal and ask it to refine the rubric for next time.

Scale it wider

Once the kit works for one role, duplicate it and ask the agent to adapt the dimensions and questions for a different role in fifteen minutes instead of starting from scratch.

Hand off to another agent

Share the completed template and rubric with the hiring manager before the debrief so they can compare their scores to yours and spot disagreements early.

Fire it from an event

Set a calendar reminder three months after someone is hired to review whether the dimensions you tested in the interview predicted their actual performance, then adjust the kit.

The work this replaces.

These are the O*NET work activities this workflow covers, and the categories of tool that address them.

Addressed by generalist agents, builders