The Discovery Question Autopsy: Which Questions Actually Advance Deals
A practical method for auditing recorded discovery calls: how to code question type and follow-ups, pick a verifiable advance as your outcome, and control confounders.
Most teams have hundreds of recorded discovery calls and a qualification framework nobody believes. The fix is not another framework debate. It is an autopsy: pick one observable deal advance from CRM field history, code every question in a sample of calls by type, object, and whether the rep followed up on the answer, then block the comparison by lead source, segment, and deal size before you believe anything. That sequence, outcome first, corpus second, coding book third, confounders fourth, is what separates a finding from a story.
I have run this on samples in the low hundreds of calls. The hardest part is never the analysis. It is resisting the urge to open transcripts before you have written down what "advanced" means.
Here is the method, including the legal design constraints most teams discover too late.
Start With the Outcome, Not the Transcript
Write the outcome definition before you open a single call. One sentence, one observable event, resolved edge cases. My default: a scheduled next meeting with a named additional stakeholder who was not on the discovery call, or a committed evaluation step with a date attached.
Pull that label from CRM field history, not from the current stage value. A stage field a rep can edit is a label a rep can retro-fit after the deal closes. Field history gives you a timestamp and an actor, which is what makes the label defensible in a pipeline review.
Why insist on a named additional stakeholder? Gartner's B2B buying research describes purchasing as a group working through parallel jobs rather than a linear path through one champion [3]. A single-threaded booked meeting is therefore a weak outcome variable. It can happen while the actual buying group remains untouched.
Resolve edge cases in writing first: does a rescheduled meeting count (yes, if it held), a no-show (no), an internal-only next step like "I'll socialize it with my team" (no, unless a named person and date appear in CRM).
| Outcome candidate | What it measures | Why it fails or holds | Verdict |
|---|---|---|---|
| Stage change | Rep's interpretation | Editable, often batch-updated at quarter end | Reject |
| Meeting booked | Calendar activity | Counts single-threaded meetings that die | Weak |
| Verbal interest in notes | Rep sentiment | Free text, unauditable, coder-dependent | Reject |
| Named-stakeholder next step | Buying-group expansion | Timestamped in field history, hard to fake | Use as primary |
| Committed evaluation step | Buyer effort with a date | Strong but rarer, thins your cells fast | Use as secondary |
Build the Corpus Without a Legal Problem
Recording law is an analysis design input, not a footnote. California Penal Code 632 restricts the recording of confidential communications without the consent of all parties to the communication [2]. If your sample mixes jurisdictions, you are sorting consent questions after the coding work is done, which is the worst possible order.
Treat all-party consent as the default standard for the entire sample. It costs you nothing analytically and removes the argument entirely. Confirm the exact notice language your recording platform plays or displays, and log which calls carried it. Calls without a documented notice are excluded from the corpus, not justified later.
If you sequence EU or UK prospects, remember that a transcript naming an individual is personal data. The GDPR gives data subjects a right of access to their personal data, so your retrieval path needs to reach transcripts, not just CRM fields [4]. UK-targeted teams should also check the ICO's direct marketing guidance for how their outreach channels are treated [5].
Then document sampling mechanics, because this is where credibility is won or lost:
- Date range with a reason (one full quarter, not "whatever was recent")
- Reps included, all of them in scope, not just the ones who record diligently
- Segments included, and the counts drawn from each
- Draw method, random within strata rather than first-available
A corpus assembled from whatever the tool happened to retain silently oversamples long calls and enterprise deals. Retention policies, storage tiers, and rep recording habits all correlate with deal size. If you skip stratified sampling, your "finding" may be nothing more than a description of your enterprise segment.
The Coding Book: Type, Object, and Follow-Up
Code three dimensions on every question. Type: open, closed, mirrored, follow-up. Object: current process, cost of inaction, decision path, budget authority. And a separate binary field for whether the rep asked a follow-up to the answer they received.
Follow-up gets its own field rather than being folded into type. Harvard Business Review's review of question-asking research distinguishes question types and identifies follow-up questions as doing distinctive conversational work [1]. In practice that means a rep can ask a strong open question, accept a one-line answer, and move on. That behavior is invisible in a keyword count and obvious in a coded transcript.
Version-control the coding book. Put the version number on every coded call. Without it, this quarter's results and next quarter's results are not comparable, and you will not be able to test whether coaching transferred.
Calibrate before you scale. Two coders label the same 25 calls independently, reconcile every disagreement, and update the coding book with the resolutions. Only then split the remaining volume. Common mis-codes to write into the book explicitly:
- Rhetorical framing coded as an open question ("So you'd want that fixed, right?")
- Stacked questions counted once instead of split into separate codes
- Confirmation restated as a follow-up ("Got it, makes sense" is not a follow-up)
- Budget authority coded when the rep asked about budget amount, which is a different object
Questions That Create the Illusion of Qualification
Budget questions asked as closed checkboxes produce an answer you can type into a field and learn nothing from. "Do you have budget for this?" gets a yes. The yes gets typed into a picklist. The deal dies in procurement six weeks later.
Authority questions that name a title invite a confident guess. Ask "are you the decision maker?" and most buyers say yes, because in their mental model they are. Ask how the last comparable purchase actually got approved and you get the chain they walked: the finance partner, the security review, the signature threshold.
BAD SET (coded: closed / budget-authority / no-follow-up)
Q: "Do you have budget allocated for this?"
Q: "Are you the decision maker on this one?"
Q: "Is this a priority for the team this year?"
BETTER SET (coded: open / decision-path / follow-up-asked)
Q: "How was the last tool in this category funded?" [open / decision-path]
Q: "Who signed that one off, and what did they ask for?" [follow-up / budget-authority]
Q: "What happened the last time this slipped a quarter?" [open / cost-of-inaction]
Q: "What did that cost the team, concretely?" [follow-up / cost-of-inaction]The pattern worth hunting in your own data is the follow-up failure: a strong question asked, a vague answer accepted. It shows up in transcripts as a good question followed immediately by the rep's next scripted question. Code it, count it, and you have a coachable behavior instead of an opinion about discovery quality. The same discipline shows up in a discovery-led demo framework, where what you demo is determined by what discovery actually surfaced.
Controlling the Confounders That Fake Your Findings
Block the comparison before you draw conclusions. Lead source, segment, and deal size are the three that reliably fake findings. Warm enterprise inbound calls get better questions and better outcomes for reasons that have nothing to do with causation.
Report raw counts inside each block. If a cell holds eight calls, print the eight and mark the pattern unresolved. A credible autopsy publishes its thin cells rather than averaging them away.
Watch for reverse causality. Some questions get asked because the deal is already going well. A rep asks about procurement timelines after the buyer volunteers that their contract expires in March. Check the transcript position: did the question precede or follow the buyer's first expansion signal? That single check kills a surprising number of candidate findings.
| Confounder | How it distorts | Control method | Cheapest check |
|---|---|---|---|
| Lead source | Inbound buyers self-qualify | Block inbound vs outbound separately | Compare advance rate by source first |
| Segment | Enterprise calls run longer | Stratify sample by segment | Count calls per segment in corpus |
| Deal size | Bigger deals get senior reps | Block by amount band | Cross-tab rep seniority by band |
| Call length | Long calls hold more questions | Normalize per-question rates | Plot question count vs duration |
| Reverse causality | Question follows buyer signal | Record question position in transcript | Flag questions after first expansion cue |
Publish an explicit unresolved list. It becomes the design spec for next quarter's sample.
Turning Coded Findings Into a Coaching Artifact Reps Use
Report at the question-behavior level, never the rep level. Behavior findings get adopted. Rep leaderboards get litigated, and the litigation consumes the meeting you wanted for coaching.
Ship one page. Three to five coded behaviors, the observed counts behind each, and the exact phrasing reps can borrow. Then retire the framework debate, because you now have evidence about which parts of your qualification model your buyers actually respond to.
Instrument the follow-through cheaply. Derive metrics from fields the sequencer and CRM already write, and reject anything requiring a new free-text field from reps. If you need a rep to type something new, the metric dies in six weeks. This is the same constraint that governs the handful of sales metrics worth maintaining.
Then re-sample. Draw a fresh set of calls next quarter using the same coding-book version and check whether the behavior transferred. That is the only proof the coaching worked. Stakeholder-coverage findings feed directly into multi-threading enterprise deals, and your outcome definition should match the one used in your SDR-to-AE handoff review so the two analyses do not contradict each other.
FAQ: Sample Size, Tooling, and What to Do With Old Calls
How many calls do I need? Enough to leave double-digit counts inside each confounder block. In practice that means starting in the several-hundred range and reporting counts per cell so readers can judge the thin ones themselves.
Can AI do the coding? Transcript-level tagging handles question type and object reasonably well. Follow-up detection and mis-code review still need human calibration on a sample, because models tend to score acknowledgments as follow-ups.
What about calls recorded before we had a documented consent notice? Exclude them from the analysis corpus. Do not retrofit a justification for a call you cannot show carried a notice.
Which tools? Conversation-intelligence platforms such as Gong, Chorus in Zoom Revenue Accelerator, and Clari Copilot all export transcripts. The analysis itself runs fine in a spreadsheet or a notebook. No new vendor required.
Does this replace MEDDPICC? No. It tells you which parts of your framework your buyers respond to and which parts exist only to populate CRM fields.
Your First Autopsy: What to Do This Week
Write the outcome definition in one paragraph and build the field-history query that produces the label. That is one working session, not a project.
Then draw the sample from calls with documented consent, tag the coding book version 1.0, and calibrate on 25 calls with a second coder before splitting the volume.
Start tracking one metric now: share of discovery calls that produced a next step with a named additional stakeholder. It is derivable from data you already have.
And calendar the re-sample date before you publish the findings. An autopsy without a follow-up sample is an opinion with a spreadsheet attached, which is exactly the thing you set out to replace.
References
[1]Harvard Business Review, "The Surprising Power of Questions," 2018. https://hbr.org/2018/05/the-surprising-power-of-questions
[2]California Legislative Information, "Penal Code Section 632." https://leginfo.legislature.ca.gov/faces/codes_displaySection.xhtml?lawCode=PEN§ionNum=632
[3]Gartner, "The B2B Buying Journey." https://www.gartner.com/en/sales/insights/b2b-buying-journey
[4]EUR-Lex, "Regulation (EU) 2016/679 (General Data Protection Regulation), consolidated text." https://eur-lex.europa.eu/eli/reg/2016/679/oj
[5]Information Commissioner's Office, "Direct marketing and privacy and electronic communications." https://ico.org.uk/for-organisations/direct-marketing-and-privacy-and-electronic-communications/
Ready to transform your sales pipeline?
See how Prospectory's AI-powered platform can help your team research, reach, and relate to prospects at scale.
Related Articles
Stop Paying SDRs Per Meeting: A Comp Redesign for the AI Era
Meeting-based SDR comp rewards booking volume that AI can increase. Use this phased framework to tie pay to held meetings, accepted opportunities, and pipeline quality.
How to Use Prospectory for Account-Based Marketing
A practical Prospectory ABM playbook for unifying target accounts, named buyers, proven signals, and sales context before orchestrating outreach.
Why AI Sales Intelligence Fails Without a Signal-to-Action System
A practical operating model for turning buyer signals into timely sales plays, coordinated stakeholder engagement, and measurable revenue outcomes across the team.