Est.

Work Sample Tests for SDR Candidates

Structured work samples predict SDR performance better than interviews alone.

Senior Writer · · 14 min read
Cover illustration for “Work Sample Tests for SDR Candidates”
Interview Design · September 19, 2026 · 14 min read · 3,216 words

SDR hiring runs on a fiction: that a good interview predicts a good rep. It doesn't, and the math behind why it doesn't should worry any sales leader still relying on conversation alone. Bridge Group's compensation research puts a fully loaded SDR seat in US B2B SaaS at somewhere between $120,000 and $160,000 a year, and a well-timed hire in that seat can influence a substantial multiple of that figure in annual pipeline. Get the hire wrong and the downside is a year of misdirected pipeline and a re-hire six months later. It's a year of misdirected pipeline and a re-hire six months later.

Gartner's turnover data makes the stakes sharper. The average SDR stays in the role about 14 months, and more than 85% leave within 18 months, with annual turnover running between 30% and 50%. RepVue's figure adds to that: only 56.7% of SDRs are hitting quota. Most seats in most companies are underperforming, and most of those seats will turn over before the underperformance even gets diagnosed properly.

The diagnosis is structural. Most interview processes test for polish and enthusiasm, because those are the easiest things to observe in fifty minutes over video. But polish and enthusiasm aren't on the list of things that actually predict SDR performance. Resilience, coachability, prospecting discipline, objection handling, and writing quality are, and none of those appear reliably in a conversation where a candidate is simply describing what they'd do. Research on this gap is blunt about the mechanism: the space between how candidates describe their approach and how they execute it is often wide, and candidates who perform well on structured work sample tests ramp faster than those who only interview well. This piece lays out a structured set of exercises, a scenario built around applying AI to the role as it is evolving, and a scoring approach that keeps evaluator bias from creeping back into the process.

What the SDR role demands in 2026, and what that means for testing

Strip the job down to its core motion: identify target accounts, build and enrich contact lists, research buyers, run multi-channel sequences across email, phone, and LinkedIn, and pass qualified meetings to account executives. An SDR's job is to create the conditions for a close, not to close. That distinction matters for testing, because it means the exercises need to isolate judgment and process, not closing charisma.

The role is also shrinking and changing shape at the same time, which makes hiring for it harder, not easier. Emergence Capital's 2026 survey of more than 560 B2B software companies found that 36% decreased SDR or BDR headcount over the past year, the highest reduction rate of any sales role surveyed, while only 19% grew their SDR teams. Agentic AI tools now absorb the high-volume, repetitive parts of prospecting: list building, initial enrichment, first-pass sequencing. What's left for the human is the judgment-intensive slice, roughly the last 20% of the workflow: tone, reading the room, sign-off on outreach before it goes out, and the relationship moments that a model can't fake convincingly yet.

Fully autonomous AI SDR deployments were tried at scale, and by early 2026 most of them had reverted to hybrid models. Human-plus-AI outperforms full automation once volume and complexity reach a meaningful threshold, which tells hiring managers something concrete: candidates now need to be evaluated on both the traditional execution skills, cold calling, objection handling, written sequencing, and the newer operator skill of directing and quality-checking AI tooling. A test suite that only covers the former is testing for last decade's job.

Five competencies produce SDR performance: resilience under rejection, coachability, prospecting discipline, objection handling, and writing quality. Of those five, coachability is the one worth weighting hardest. MeetMiki's 2026 research on SDR ramp speed identifies coachability as the single strongest predictor of how fast a new hire gets productive, for a reason that's almost too obvious to state directly: technique is teachable. Willingness to be taught isn't, at least not on a hiring manager's timeline.

The four canonical work sample exercises and what each one reveals

Diagram: Trigger Events vs. Cold Outreach: The Win Rate Gap. Visualizes: Show a stark magnitude contrast between two win rates drawn directly from Champify's 2025 research cited in the article: outreach into accounts with an active buying trigger…

These four exercises work as a coordinated sequence. Later exercises build on signals the earlier ones reveal, so skipping around undercuts the whole design.

Exercise 1: the async video screen. Before any live interaction, ask the candidate to record a two-to-three-minute Loom-style video introducing themselves and describing a recent SDR win. This runs before the live process starts, and it screens for something simple: can this person communicate with clarity and presence when there's no interviewer nodding along to smooth over the awkward parts? Evaluate understandability, poise, and above all the specificity of the "win" they describe. A vague win, "I crushed quota" with no numbers or mechanism attached, signals a pattern of describing activity as if it were outcome. Talent Hackers' 2025 guidance on this exercise is to weight clarity and whether the video sounds like a human talking to a human over production polish. Keep this stage under an hour of total candidate time. It's fair, and it's legally defensible.

Exercise 2: the cold call role play. This is where the process gets live. The hiring manager plays a skeptical prospect, a VP of Sales or equivalent, and the candidate places the cold call cold. What surfaces here is the candidate's ability to earn the right to keep talking, their instinct for objection handling under real pressure, and whether they can think on their feet while still following some kind of structure. One strong tell: candidates who ask at least two qualifying questions before pitching anything. That single behavior separates reps who work to earn the conversation from reps who launch straight into the deck regardless of what the prospect actually said.

The real signal comes after the first run. Give the candidate one specific, pointed piece of feedback, "your opening was too long" works well, and have them run the call again immediately. What's being measured is whether they process feedback and adjust in real time. It's whether they process feedback and adjust in real time, or whether they get defensive and repeat roughly the same call. MeetMiki's 2026 research argues this single live test is worth more than three hypothetical interview questions combined, because it reveals in real time whether a candidate can absorb coaching and adjust their execution rather than repeat the same approach. For context when grading: Cold call dial-to-meeting conversion runs around 2.3% on average, with top performers reaching 5% to 8% through discipline rather than raw volume. The exercise is looking for that discipline in embryo.

Exercise 3: prospect research and sequence build. Give the candidate a fictional prospect profile: name, title, company, a short ICP summary, and one stated trigger event, recent funding, a leadership change, a product launch. Ask them to build a six-step outreach sequence, specifying channel, message purpose, and timing for each touch. This reveals research depth, whether their targeting actually fits the stated ICP, whether they default to generic volume tactics or persona-specific messaging, and critically, whether they use the trigger event.

The trigger event determines outcome: Champify's 2025 research found that selling into accounts with an active buying trigger produces a 37% win rate, against 19% for cold outreach with no trigger. Champify's 2025 research found that selling into accounts with an active buying trigger produces a 37% win rate, against 19% for cold outreach with no trigger. A candidate who ignores the one piece of signal handed to them on a plate is telling the grader something important about how they'll operate with real accounts. On cadence, the general benchmark for mid-market B2B sequences runs 10 to 12 touches across four to six weeks, with enterprise sequences stretching to 12 to 18 touches across eight to twelve weeks. A candidate's instinct on density, do they front-load LinkedIn and email, introduce phone from touch four onward, include an actual breakup message, is itself a data point about their judgment, separate from whether the exact numbers match the benchmark. Messaging that shifts by persona across the sequence is a good sign. Messaging that reads the same for every touch, regardless of channel or stage, tells the grader the candidate defaults to volume over relevance.

Exercise 4: the written email sequence. Using the same fictional prospect, ask for three emails: initial outreach, a follow-up, and a final breakup attempt. This isolates writing quality, tone control, and the candidate's grasp of a subtle point that a lot of reps never learn, each email in a sequence has a distinct job, and treating all three the same way is a tell. Strong candidates write subject lines that earn opens, open with a first line referencing something specific about the prospect rather than a template phrase, and close with a CTA that makes the next step obvious and low-friction. Weak candidates lean on "I hope this finds you well," dump product features into email one, and write the breakup message in the same tone as the introduction.

Grading here should account for where the industry actually sits. Average B2B cold email reply rates now run around 3.43% to 5.1%, down from roughly 8.5% in 2019, driven by inbox saturation and a flood of low-effort AI-generated outreach. The exercise is calibrated to find candidates who write meaningfully above that noise floor, not candidates who write competently by 2019 standards. Expandi's figures put LinkedIn DMs reply rates at around 10.3%, against 5.1% for cold email, and this belongs in the grading conversation. A candidate who proposed LinkedIn as the actual first touch back in Exercise 3, then wrote a strong email for touch two here in Exercise 4, is demonstrating channel awareness on top of copywriting skill, and that combination is worth more than either skill alone.

Adding an AI-operator scenario to test for the skill that matters in 2026

The SDR job is shifting from executing sequences to operating systems, and hiring processes that don't test for that shift are testing for a version of the role that's already fading. A fifth scenario closes that gap directly.

Give the candidate an AI-drafted cold email, or an AI-generated prospect list, and ask them to do three things: evaluate its quality, identify what's wrong or missing, and describe how they'd adjust or reprompt the tool to produce something they'd actually be willing to send. What this reveals has nothing to do with whether the candidate can operate a particular piece of software. It's about whether they treat AI output as a first draft that requires judgment, or as a finished product ready to ship as-is. A candidate who accepts a generic, unpersonalized AI-generated email without flagging the ICP mismatch or the missing trigger event is showing the grader how they'll treat real AI tooling once they're inside a live pipeline.

The strongest candidates will articulate, without being prompted for it directly, a version of the guardrail that's becoming standard across GTM teams: a human owns every customer-facing output before it goes out, AI suggestions function as drafts and inputs rather than decisions, and the data any model runs on needs to be governed, not assumed clean. Teams getting this right use agentic tools to eliminate the repetitive 80% of the workflow, list building, first-pass enrichment, initial drafting, while keeping a human on the 20% that requires reading a room or making a judgment call a model can't make responsibly. A candidate who can describe that division of labor unprompted is describing the actual job as it exists now, not the job as it existed five years ago.

This scenario doesn't need its own hour. It folds cleanly into Exercise 3, ask the candidate to critique an AI-drafted version of the sequence before building their own, or it runs as a brief standalone add-on. Either way, it adds a layer of signal without adding meaningful time to the process. Platforms like Apollo, which combine account and contact data, buying-signal detection, and AI-assisted drafting into one connected system, are close to the actual environment a hired SDR will work inside day to day. Candidates who can talk about using a connected system like that as a co-pilot, something that removes repetitive load while leaving judgment calls to a human, are interviewing for the job as it's actually structured in 2026, not for a job that stopped existing.

How to score the exercises without letting bias back in

Work sample tests only solve the polish bias if the scoring process doesn't quietly reintroduce it. Without a rubric, evaluators default to a gut read, "would I want to work with this person?", and that question is a proxy for exactly the charisma and likability bias the exercises were built to route around.

Score each exercise against the same five competencies every time: resilience and rejection recovery, coachability, prospecting discipline, objection handling, writing quality. Not overall impression. A one-to-five scale per competency per exercise gives a grid rather than a gut feeling, and MeetMiki's 2026 scorecard approach uses a one-to-five scale per competency per exercise, a structure that surfaces consistent signal rather than letting one strong moment, a great opening line, a smooth recovery on the second call attempt, carry the whole evaluation. A candidate who's strong in one exercise and weak in three others should score as weak overall, and a grid makes that visible in a way a single "great candidate" impression doesn't.

Research quality deserves tracking as a thread that runs across exercises 2, 3, and 4, rather than being scored only in the exercise explicitly labeled "research." Did the candidate show evidence of understanding the persona before writing the emails? Before preparing for the call? Consistency, or its absence, across those three touchpoints is more diagnostic than strong performance in any single one.

Coachability scoring in the role play has a clean binary at its center: strong candidates pause, visibly process the feedback, and deliver a tighter version of the call. Weak candidates get defensive, get flustered, or essentially repeat the same call with cosmetic changes. Objection handling has a similar binary: strong candidates stay in the conversation when the prospect tries to end it, and use some recognizable structure, acknowledge, reframe, advance. Weak candidates capitulate immediately, pivot to "I'll just send you an email," or bury the prospect in over-explanation. Writing scores should anchor on specificity of the first line, subject line quality, CTA clarity, and tone variation across the three emails, not on grammar alone; grammar is the least informative thing about a cold email.

Two logistics guardrails matter as much as the rubric itself. Keep unpaid exercises under one hour of total candidate time, and compensate for any exercise that runs longer, both because it's the right thing to do and because it signals to candidates how the company actually operates before they've accepted an offer. And have multiple evaluators score independently before comparing notes. Calibration sessions afterward are where ambiguous rubric language gets caught, usually because two evaluators scored the same performance three points apart and neither can explain why without a conversation.

The interview questions that wrap around the exercises to complete the picture

Exercises show what a candidate can actually do. Targeted questions show why they make the choices they make, and how they respond when a conversation turns uncomfortable, which the exercises alone don't always capture.

MeetMiki's 2026 framework offers a clean coachability question: "Tell me about a time you got direct, critical feedback from a manager. What was it, and what did you do with it?" A strong answer names the specific feedback, describes a specific behavior change, and points to a measurable result. A weak answer stays at "I'm always open to feedback" with nothing behind it. Pair that with a live version of the same test: mid-interview, tell the candidate their last answer ran too long, and ask them to redo it. The goal is watching whether they adjust or defend, same mechanism as the cold call re-run, different context.

On prospecting discipline, two questions do most of the work. "Walk me through how you research a prospect before your first outreach, give me a real example," separates candidates who name specific sources and specific trigger events from candidates who say "I look at their LinkedIn and then call." And "how do you decide who to call first when you sit down in the morning?" separates candidates with an actual prioritization system, follow-ups first, then the trigger-event list, then cold accounts, from candidates who simply say "I just start dialing."

Resilience gets tested with a question built around a bad day rather than a good one: "Describe a day when nothing went right, zero meetings, an inbox full of not-interested replies. What did you do at 4 PM?" A strong answer names a plan, reviewing messaging, adjusting approach, auditing the list. A weak answer says "I stayed positive," which is a feeling, not a strategy, and doesn't tell the grader anything about behavior.

Eight to twelve questions across these competency categories, plus the role-play prompt, covers the ground without repetition; MeetMiki's 2026 research notes that adding more questions past that range produces diminishing signal, not additional insight. The most useful move an evaluator can make is connecting interview answers back to exercise performance directly: if a candidate claims disciplined research habits in the interview but their Exercise 3 sequence never touched the stated trigger event, that gap should be named and probed, not smoothed over.

What a full SDR hiring process looks like end to end

Put together, the process runs in roughly three to four steps and stays under two hours of total candidate time, which matters both for candidate experience and for keeping the funnel from bleeding good candidates who have other offers moving faster.

Step one is an async screen: a resume and LinkedIn scan for ICP alignment, outbound versus inbound experience, and familiarity with the relevant tools, CRM platforms, sales engagement tools, a prospecting add-on for a professional networking platform, paired with the two-to-three-minute Loom video and a short cold email sample against a fictional brief. Step two is a screening call, twenty to thirty minutes, covering motivation for the role, product and market awareness, and a surface read on prospecting experience. This call is calibration. It's where the hiring manager decides whether the candidate is worth the time investment of the live exercises, not where the hiring decision actually gets made.

Step three is the work sample session itself: the cold call role play with its real-time coachability re-run, the prospect research and sequence build, and the written three-email sequence, with the AI-operator scenario either folded into the research exercise or run as its own ten-minute add-on. Step four closes the loop: independent scoring against the five-competency rubric from every evaluator involved, a calibration conversation to reconcile any scores that diverge sharply, and a final decision built from the scoring grid rather than from whoever in the room made the strongest personal impression.

None of this removes judgment from hiring. It relocates judgment to where it belongs, on structured evidence of how a candidate actually performs under the specific pressures of the job, rather than on how comfortable they made an interviewer feel for fifty minutes. Given what a bad SDR hire costs against the pipeline a good one generates, that relocation is not optional. It's the job.

Sources

  1. How to Evaluate Offshore SDR Candidates in 2025
  2. How to Hire an SDR in 2026 (Complete Employer Guide) - RevPilots
  3. Should You Hire an SDR? The 4 Tests That Decide
  4. 25 SDR Interview Questions Hiring Managers Should Ask (+ Scorecard)
  5. B2B Companies Hiring SDRs, BDRs, and ADRs: What the 2026 Hiring Data Reveals
  6. lathire.com
Filed underInterview Design

More in Interview Design