Sales Strategy

The Discovery Question Autopsy: Which Questions Actually Advance Deals

A practical method for auditing recorded discovery calls: how to code question type and follow-ups, pick a verifiable advance as your outcome, and control confounders.

SA
Simone Adeyemi
Head of Growth Intelligence
September 28, 20269 min

Most teams have hundreds of recorded discovery calls and a qualification framework nobody believes. The fix is not another framework debate. It is an autopsy: pick one observable deal advance from CRM field history, code every question in a sample of calls by type, object, and whether the rep followed up on the answer, then block the comparison by lead source, segment, and deal size before you believe anything. That sequence, outcome first, corpus second, coding book third, confounders fourth, is what separates a finding from a story.

I have run this on samples in the low hundreds of calls. The hardest part is never the analysis. It is resisting the urge to open transcripts before you have written down what "advanced" means.

Here is the method, including the legal design constraints most teams discover too late.

Start With the Outcome, Not the Transcript

Write the outcome definition before you open a single call. One sentence, one observable event, resolved edge cases. My default: a scheduled next meeting with a named additional stakeholder who was not on the discovery call, or a committed evaluation step with a date attached.

Pull that label from CRM field history, not from the current stage value. A stage field a rep can edit is a label a rep can retro-fit after the deal closes. Field history gives you a timestamp and an actor, which is what makes the label defensible in a pipeline review.

Why insist on a named additional stakeholder? Gartner's B2B buying research describes purchasing as a group working through parallel jobs rather than a linear path through one champion [3]. A single-threaded booked meeting is therefore a weak outcome variable. It can happen while the actual buying group remains untouched.

Resolve edge cases in writing first: does a rescheduled meeting count (yes, if it held), a no-show (no), an internal-only next step like "I'll socialize it with my team" (no, unless a named person and date appear in CRM).

Outcome candidateWhat it measuresWhy it fails or holdsVerdict
Stage changeRep's interpretationEditable, often batch-updated at quarter endReject
Meeting bookedCalendar activityCounts single-threaded meetings that dieWeak
Verbal interest in notesRep sentimentFree text, unauditable, coder-dependentReject
Named-stakeholder next stepBuying-group expansionTimestamped in field history, hard to fakeUse as primary
Committed evaluation stepBuyer effort with a dateStrong but rarer, thins your cells fastUse as secondary

Recording law is an analysis design input, not a footnote. California Penal Code 632 restricts the recording of confidential communications without the consent of all parties to the communication [2]. If your sample mixes jurisdictions, you are sorting consent questions after the coding work is done, which is the worst possible order.

Treat all-party consent as the default standard for the entire sample. It costs you nothing analytically and removes the argument entirely. Confirm the exact notice language your recording platform plays or displays, and log which calls carried it. Calls without a documented notice are excluded from the corpus, not justified later.

If you sequence EU or UK prospects, remember that a transcript naming an individual is personal data. The GDPR gives data subjects a right of access to their personal data, so your retrieval path needs to reach transcripts, not just CRM fields [4]. UK-targeted teams should also check the ICO's direct marketing guidance for how their outreach channels are treated [5].

Then document sampling mechanics, because this is where credibility is won or lost:

  • Date range with a reason (one full quarter, not "whatever was recent")
  • Reps included, all of them in scope, not just the ones who record diligently
  • Segments included, and the counts drawn from each
  • Draw method, random within strata rather than first-available
The Most Common Autopsy Failure

A corpus assembled from whatever the tool happened to retain silently oversamples long calls and enterprise deals. Retention policies, storage tiers, and rep recording habits all correlate with deal size. If you skip stratified sampling, your "finding" may be nothing more than a description of your enterprise segment.

The Coding Book: Type, Object, and Follow-Up

Code three dimensions on every question. Type: open, closed, mirrored, follow-up. Object: current process, cost of inaction, decision path, budget authority. And a separate binary field for whether the rep asked a follow-up to the answer they received.

Follow-up gets its own field rather than being folded into type. Harvard Business Review's review of question-asking research distinguishes question types and identifies follow-up questions as doing distinctive conversational work [1]. In practice that means a rep can ask a strong open question, accept a one-line answer, and move on. That behavior is invisible in a keyword count and obvious in a coded transcript.

Version-control the coding book. Put the version number on every coded call. Without it, this quarter's results and next quarter's results are not comparable, and you will not be able to test whether coaching transferred.

Calibrate before you scale. Two coders label the same 25 calls independently, reconcile every disagreement, and update the coding book with the resolutions. Only then split the remaining volume. Common mis-codes to write into the book explicitly:

  • Rhetorical framing coded as an open question ("So you'd want that fixed, right?")
  • Stacked questions counted once instead of split into separate codes
  • Confirmation restated as a follow-up ("Got it, makes sense" is not a follow-up)
  • Budget authority coded when the rep asked about budget amount, which is a different object

Questions That Create the Illusion of Qualification

Budget questions asked as closed checkboxes produce an answer you can type into a field and learn nothing from. "Do you have budget for this?" gets a yes. The yes gets typed into a picklist. The deal dies in procurement six weeks later.

Authority questions that name a title invite a confident guess. Ask "are you the decision maker?" and most buyers say yes, because in their mental model they are. Ask how the last comparable purchase actually got approved and you get the chain they walked: the finance partner, the security review, the signature threshold.

BAD SET (coded: closed / budget-authority / no-follow-up)
Q: "Do you have budget allocated for this?"
Q: "Are you the decision maker on this one?"
Q: "Is this a priority for the team this year?"

BETTER SET (coded: open / decision-path / follow-up-asked)
Q: "How was the last tool in this category funded?"      [open / decision-path]
Q: "Who signed that one off, and what did they ask for?" [follow-up / budget-authority]
Q: "What happened the last time this slipped a quarter?" [open / cost-of-inaction]
Q: "What did that cost the team, concretely?"            [follow-up / cost-of-inaction]

The pattern worth hunting in your own data is the follow-up failure: a strong question asked, a vague answer accepted. It shows up in transcripts as a good question followed immediately by the rep's next scripted question. Code it, count it, and you have a coachable behavior instead of an opinion about discovery quality. The same discipline shows up in a discovery-led demo framework, where what you demo is determined by what discovery actually surfaced.

Controlling the Confounders That Fake Your Findings

Block the comparison before you draw conclusions. Lead source, segment, and deal size are the three that reliably fake findings. Warm enterprise inbound calls get better questions and better outcomes for reasons that have nothing to do with causation.

Report raw counts inside each block. If a cell holds eight calls, print the eight and mark the pattern unresolved. A credible autopsy publishes its thin cells rather than averaging them away.

Watch for reverse causality. Some questions get asked because the deal is already going well. A rep asks about procurement timelines after the buyer volunteers that their contract expires in March. Check the transcript position: did the question precede or follow the buyer's first expansion signal? That single check kills a surprising number of candidate findings.

ConfounderHow it distortsControl methodCheapest check
Lead sourceInbound buyers self-qualifyBlock inbound vs outbound separatelyCompare advance rate by source first
SegmentEnterprise calls run longerStratify sample by segmentCount calls per segment in corpus
Deal sizeBigger deals get senior repsBlock by amount bandCross-tab rep seniority by band
Call lengthLong calls hold more questionsNormalize per-question ratesPlot question count vs duration
Reverse causalityQuestion follows buyer signalRecord question position in transcriptFlag questions after first expansion cue

Publish an explicit unresolved list. It becomes the design spec for next quarter's sample.

Turning Coded Findings Into a Coaching Artifact Reps Use

Report at the question-behavior level, never the rep level. Behavior findings get adopted. Rep leaderboards get litigated, and the litigation consumes the meeting you wanted for coaching.

Ship one page. Three to five coded behaviors, the observed counts behind each, and the exact phrasing reps can borrow. Then retire the framework debate, because you now have evidence about which parts of your qualification model your buyers actually respond to.

Instrument the follow-through cheaply. Derive metrics from fields the sequencer and CRM already write, and reject anything requiring a new free-text field from reps. If you need a rep to type something new, the metric dies in six weeks. This is the same constraint that governs the handful of sales metrics worth maintaining.

Then re-sample. Draw a fresh set of calls next quarter using the same coding-book version and check whether the behavior transferred. That is the only proof the coaching worked. Stakeholder-coverage findings feed directly into multi-threading enterprise deals, and your outcome definition should match the one used in your SDR-to-AE handoff review so the two analyses do not contradict each other.

FAQ: Sample Size, Tooling, and What to Do With Old Calls

How many calls do I need? Enough to leave double-digit counts inside each confounder block. In practice that means starting in the several-hundred range and reporting counts per cell so readers can judge the thin ones themselves.

Can AI do the coding? Transcript-level tagging handles question type and object reasonably well. Follow-up detection and mis-code review still need human calibration on a sample, because models tend to score acknowledgments as follow-ups.

What about calls recorded before we had a documented consent notice? Exclude them from the analysis corpus. Do not retrofit a justification for a call you cannot show carried a notice.

Which tools? Conversation-intelligence platforms such as Gong, Chorus in Zoom Revenue Accelerator, and Clari Copilot all export transcripts. The analysis itself runs fine in a spreadsheet or a notebook. No new vendor required.

Does this replace MEDDPICC? No. It tells you which parts of your framework your buyers respond to and which parts exist only to populate CRM fields.

Your First Autopsy: What to Do This Week

Write the outcome definition in one paragraph and build the field-history query that produces the label. That is one working session, not a project.

Then draw the sample from calls with documented consent, tag the coding book version 1.0, and calibrate on 25 calls with a second coder before splitting the volume.

Start tracking one metric now: share of discovery calls that produced a next step with a named additional stakeholder. It is derivable from data you already have.

And calendar the re-sample date before you publish the findings. An autopsy without a follow-up sample is an opinion with a spreadsheet attached, which is exactly the thing you set out to replace.

References

[1]Harvard Business Review, "The Surprising Power of Questions," 2018. https://hbr.org/2018/05/the-surprising-power-of-questions

[2]California Legislative Information, "Penal Code Section 632." https://leginfo.legislature.ca.gov/faces/codes_displaySection.xhtml?lawCode=PEN&sectionNum=632

[3]Gartner, "The B2B Buying Journey." https://www.gartner.com/en/sales/insights/b2b-buying-journey

[4]EUR-Lex, "Regulation (EU) 2016/679 (General Data Protection Regulation), consolidated text." https://eur-lex.europa.eu/eli/reg/2016/679/oj

[5]Information Commissioner's Office, "Direct marketing and privacy and electronic communications." https://ico.org.uk/for-organisations/direct-marketing-and-privacy-and-electronic-communications/

S

Simone Adeyemi

Head of Growth Intelligence

A contributor to Prospectory's practical guides for modern go-to-market teams.