tl;dr of Behavioral Interview Questions

  • Behavioral interview questions get you evidence. STAR is the frame for eliciting it, not the thing you score.
  • Score how strongly the example demonstrates a competency you defined before the interview, against written anchors.
  • Anchors describe observable behavior, not adjectives. Add an insufficient evidence code for competencies the interview never established.
  • Score immediately and independently, then average the numbers. Use the debrief to surface missed evidence, not to renegotiate scores.
  • Cap follow-up probes at two, so one candidate does not get more help than the next.
  • You cannot reliably spot a rehearsed answer. Ask for checkable details instead.
  • When two interviewers disagree, the anchor language settles it.
Оглавление

Behavioral interview questions ask a candidate to describe a situation in which they actually did something with results, rather than speaking about it hypothetically. There is a particular framework that interviewers can use called STAR, which stands for:

  • Situation
  • Задача
  • Действие
  • Результаты

A lot of information out there is geared toward how an interviewee can manage these questions, how to prepare, and what to consider when answering. But if you are new to interviewing, or want to get better at it, this is how you, as the interviewer, can use STAR in your interviews, how to score them, and how to build a workflow that makes sure you are hitting the right points to make the interviews as thorough and useful as possible.

The key thing to remember when scoring a STAR interview is you are not scoring the story, but the evidence that it actually reveals.

STAR is laid out in a way that makes sure you have enough information to make a solid call. But if you score purely on how the acronym is laid out, you will end up scoring the story. If you are interviewing salespeople, that can be great. For another role it does not necessarily give you what you need to make the right decision.

How to Score a Behavioral Interview Answer

The key thing to remember when scoring a behavioral interview answer is that you should have defined a competency before the interview itself. You then score against your written anchors straight away after the answer, before conferring with anybody else. An anchor is a description of what a score looks like in behavior, so “explained a technical choice to a non-technical manager and changed their position” rather than “is a good communicator”.

To make this easier, here is the process that should help you score consistently and accurately.

  1. Define the competency first. Decide the behavior that the role you are hiring for actually needs before you write down what “good” looks like. A support lead role might need the ability to de-escalate under pressure. For an account executive, it could be about walking away from a negotiation that was never going to close. If you cannot identify something before the interview, you will end up making the story fit the narrative afterward, based on your own bias.
  2. Take factual notes under each STAR element. Write down what the candidate said rather than your feeling towards it. For example, “Reworked the escalation process after missed SLAs” rather than “Strong”.
  3. Compare the evidence against the written anchors. From what the candidate has said, work out the level of evidence it reaches. Do this before any gut choices.
  4. Score immediately, with a rationale tied to a specific thing that was said. Having a solid record is key, and the longer it is left the harder it is to recall.
  5. Do not discuss the score with anyone until every interviewer has committed theirs.

The biggest positive flag in this is action. Trade-offs, decisions, and what a candidate chose not to do are the evidence itself. The additional context of the situation and the specific task lets you rank that evidence against how difficult it was. Two candidates could describe a similar action, but the context around it shows you who is stronger.

Missing sections, or anchors that are not met across all questions, should not be an automatic fail. An answer can contain all of the STAR elements and still score low, because complete is not the same as strong.

Score these things:

  • Relevance of the example
  • Personal ownership and accountability
  • Difficulty of the situation
  • Quality of reasoning
  • Evidence that the action produced something

Pay attention to, but do not score:

  • Уверенность
  • Likeability
  • Eye contact
  • Whether they said “I” a lot
  • Fluency


These can all be coached and prepped for, unlike previous results.

Score the Competency, Not the Four Letters

There are two scoring guides typically mentioned when it comes to STAR, and they measure different things.

Per-competency scoring

This gives one score per competency, with the evidence taken from anywhere in the answer. So if the candidate tells the story in a slightly non-narrative way, it can still score high, as long as reasoning and ownership are there.

Per-element scoring

This weighs each of the letters separately, then puts them together for an overall score. The drawback is that it rewards narrative competence, meaning the candidate who has practiced what they are saying, rather than the person who hit all the markers in a mixed-up order. It still has its place. If the role you are filling needs narrative construction, such as sales or marketing, it is a useful way to score.

Building a Rating Scale With Behavioral Anchors

A behavioral rating scale uses anchors written before the interview, so vague labels like “excellent communicator” get swapped for descriptions of observable behavior.

Here is a five-point scale that works across most competencies as a starting point.

ScoreLabelWhat the evidence looks like
5СильныйDifficult situation with real constraints. Candidate owned the decision, can describe what they weighed and rejected, and states an outcome they can point to.
4ТвёрдыйClear personal ownership. Actions described specifically. Reasoning present but thinner, or the situation was less demanding. Outcome stated.
3AdequateReal, relevant example. Personal contribution is identifiable. Actions described at a general level with limited reasoning.
2ОграниченныйExample is real, but the candidate’s own contribution is unclear, or the actions described belong to the team rather than to them.
1PoorNo specific example after probing, a hypothetical instead of a real one, or an example that demonstrates the opposite of the competency.
IEInsufficient evidenceThe interview did not establish enough to rate this competency. Not a low score. A gap in the interview.

The IE is important here because sometimes, when time is tight or you forgot to probe during the interview, a score can end up corrupted because people mark it low rather than actually having gained the answer.

The scale above works as a starting point for any competency, which is also its weakness. Two interviewers reading “clear personal ownership” can still picture different answers.

This is when having question-specific anchors: written examples of what a 5 and a 2 actually sound like for that exact question, in your organization. Interviewers stop interpreting the anchor and start matching against a real answer.

Making Scores Comparable Across Interviewers

In any process there will likely be multiple interviewers, which makes decisions fairer by bringing in different viewpoints and different stakeholders. Despite best efforts, people bring their own biases into interview situations. That is unavoidable. But a solid control baseline allows for a much more democratic, evidence-based approach to hiring.


Three things make interview scores comparable:

  • Identical questions
  • Independent scoring
  • Combining the numbers before discussion
  1. Standardize the questions, the order and the time. Conway, Jako and Goodman’s meta-analysis of 111 interrater reliability coefficients found that standardizing questions raised reliability more when interviewers saw candidates separately than when a panel watched the same answer. So if your interviewers run their own calls, standardization is doing more work for you, not less. The same analysis put the ceiling on validity at .67 for highly structured interviews against .34 for unstructured ones.
  2. Score independently, before anyone speaks. This makes sure no bias creeps into other people’s scores. If someone more senior scores a candidate higher and says so out loud, there is a good chance it moves everyone else’s scoring.
  3. Combine the scores mechanically, not by discussion. Figures are figures, and they do not hold strong qualitative opinions. Average the independent scores, then use the debrief to discuss where people felt evidence was missed.
  4. Hold everything else as a control. Keep the same panel for all candidates where you can, and compare each candidate to the original anchors rather than against each other.

Follow-Up Questions for Incomplete or Hypothetical Answers

STAR rewards candidates for providing real-life examples, but sometimes a candidate will not have anything to hand, or will give a hypothetical answer. That is your chance to follow up and probe deeper. Two probes is the sweet spot before scoring and moving on. A solid cap makes it fairer, and stops a likeable candidate getting more chances to expand than the next person.

What is missingProbe
Situation“What made this harder than a normal week?”
Ownership“What part of that was yours specifically?”
Действие“Walk me through the first thing you did after that.”
Рассуждения“What else did you consider, and why did you rule it out?”
Результат“How did you know it had worked?”
Learning“What would you do differently now?”
Wall-to-wall “we”“Take me to a moment where you personally decided something.”

If the candidate gives a hypothetical answer, redirect once and be specific: “Can you give me a time this actually happened, even if it was small.” If the second answer is also hypothetical, score it accordingly.

One thing to keep in mind with follow-ups. A study found that follow-up questioning can increase faking behavior, so probe to retrieve missing evidence rather than to give someone a second run at impressing you.

How to Tell a Rehearsed Answer From a Real One

The short answer is you can’t. A meta-analysis of deception judgments covered 206 studies and 24,483 judges. People average 54% correct on lie-truth judgments, and professionals do no better than anyone else. Audible lies are caught more often than visible ones, which runs counter to every body-language module ever sold to a hiring team.

So the idea that someone is “too polished” does not fly.

The issue is that preparation and fabrication are different things. Levashina and Campion’s Interview Faking Behavior scale, built across six studies with 1,346 participants, separates ordinary preparation from outright invention. Embellishing and tailoring sit at one end. Inventing and borrowing someone else’s work sit at the other.

In that study, more than 90% of candidates faked in some form.

Practicing for an interview is compliance, not deception. Companies routinely ask candidates to prepare examples beforehand, so inventing is the only type of faking that should be marked down.

Common Behavioral Interview Questions

For examples of what kind of questions are useful in behavioral interviews, here are eight, broken down by competency, with the probe that matters most for each one and what weak and strong answers sound like.

ВопросКомпетентностьBest probeWeak evidenceStrong evidence
“Tell me about a time you disagreed with a decision your manager made.”Constructive challenge“What did you do after the decision stood?”Describes being right and being ignored.Names the argument they made, the evidence behind it, and what they did once it was settled.
“Describe a time you had to deliver bad news to a customer or stakeholder.”Difficult communication“What exactly did you say first?”Talks about the situation, not the conversation.Reconstructs the actual words, the reaction, and the follow-up.
“Tell me about something you shipped that did not work.”Ownership and learning“What did you change about how you work?”Blames scope, timeline or another team.Names their own contribution to the failure and a specific behavior change since.
“Walk me through a time you had to prioritize with more work than time.”Judgment under constraint“What did you drop, and who did you tell?”Says they worked late and got it all done.Describes an explicit trade-off, the reasoning, and who they communicated it to.
“Describe a time you changed your mind based on data.”Intellectual honesty“What did you believe before?”Vague conversion story with no prior position.States the original position, the specific evidence, and the cost of switching.
“Tell me about a time you had to work with someone difficult.”Сотрудничество“What did you try that did not work?”Character assassination with no action.Describes what they adjusted in their own behavior first.
“Give me an example of a process you improved.”Initiative“Who asked you to do it?”Describes a process they were assigned to fix.Unprompted, with a before and after they can quantify or describe.
“Tell me about a time priorities changed on you mid-project.”Adapting to change“What did you have to unwind?”Frames it as somebody else’s mistake.Describes what they unwound and who they told, without turning it into a grievance.

Ask the same eight of everyone in the same order, and do not let a good answer to question three change how you hear question five.

When Two Interviewers Disagree

When two interviewers come out with a different score on the same answer, the anchor language is what settles it.

Take a hypothetical. A candidate describes a busy quarter with multiple overlapping launches. They moved a deadline, told the stakeholder before the slip rather than afterwards, but could not give a quantifiable metric for the impact because that was owned by another team.

Interviewer one scored it a 4. Interviewer two scored it a 2.

Interviewer two’s reasoning: “There’s no outcome, so there is no proof that the choice made was right.”

Interviewer one argues: “But they gave clear personal ownership, they took specific actions and they stated an outcome.” Interviewer two is mixing up “stated outcome” and “measured outcome”.

In this situation, the candidate made a decision, communicated it before things went wrong, and the number they did not have was not theirs to own.

Interviewer one is correct, by the anchor wording. Interviewer two is arguing with the rubric rather than about the candidate.

What to Record So the Panel Has Something to Calibrate Against

In the example above, the interviewers were relying on memory. Not having a clear, verbatim transcript of how the call went leads to misremembering and heated debriefs. When you have two sets of notes taken by two people who disagree, neither is a reliable resource to lean on when scoring.

By using an AI meeting assistant such as tl;dv, you record every interview and the panel gets a true transcript to refer back to. The transcript settles what was said. The anchor then settles what it was worth, and it becomes a case of making a decision rather than arguing about the small details.

It also means the interviewers are present and paying attention, rather than scratching down handwritten notes between questions. You ask, you listen, and you probe when the evidence is thin.

You can also create your own meeting templates in tl;dv, so your behavioral interview questions come with the prompts and scoring structure already built in. Name your competencies as the sections, ask for the candidate’s own words under each one, and the notes come back mapped onto your scorecard rather than onto a generic hiring summary. The automation rules will then apply it to anything with “interview” in the meeting title.

Recording gives you the record, and it comes with an obligation. Consent rules differ across US states and again across the UK and EU, so check what the rules actually require and tell candidates at the invite stage rather than at the top of the call. Telling them early tends to help you anyway. A candidate who knows the conversation is on the record also knows their examples are checkable, and that is the single thing most likely to keep an answer honest.

You can also connect the tl;dv MCP server, available on the Pro plan, and use Claude or ChatGPT with read-only access. Paste your anchor language in, query it against the transcript, and get a second read to help with scoring. Treat that output as support rather than as a score. A model reading a transcript has far less context about the role than you do, but it gives you another datapoint. The scoring decision stays with the panel, and the reasoning you write down should be a person’s.

Try it for yourself in your next round of interviews by checking out tl;dv today.

FAQs About Behavioral Interview Questions

Compare the evidence against written behavioral anchors for the competency, score immediately and independently, and record a rationale tied to something the candidate said.

Four to six competencies in a 45-minute interview, one question each with room to probe. More questions means less evidence per competency, not more coverage.

Behavioral questions ask what a candidate did. Situational questions ask what they would do. Research on interview faking found situational questions produced more faking than past-behavior ones.

Not reliably. People average 54% accuracy on lie-truth judgments and professionals do no better. Ask for checkable details and verify them afterward.