esc
§ 07
Part 7 of 12

JosephB — Sagetap Call Debrief

Campaign: Engineering Leaders: See What Your AI Coding Agents Are Actually Doing Platform: Sagetap (Market Awareness, 1,250 credits) Date: 2026-05-28 Duration: ~32 minutes Participants: Bryan (CEO), Paolo (product advisor), JosephB Outcome: Passed on Sagetap. “Tool feels not mature enough in terms of providing end-to-end value.”

Pre-call evaluation: 2026-05-28-sagetap-josephb-evaluation Pre-call briefing: 2026-05-28-josephb-meeting-prep


Sagetap Feedback

CategoryRating
Overall3.6 / 5
Differentiation4 / 5
Value3 / 5
Timing4 / 5
Integration4 / 5
Cost3 / 5
Presenter3 / 5

Status: Passed

His note: “Prefer not to share my details at the moment, until I decide if I want to proceed with this vendor or not.”

His product feedback: “Tool feels not mature enough in terms of providing end-to-end value. Feels like it just delivers information, but without specific insights, suggested action items, or downstream process.”

Read: 60% polite no, 40% genuine maybe. He saw the market need (4/5 differentiation, 4/5 timing) but didn’t see the product delivering specific value (3/5 value). The VP mention could be a brush-off or a real referral. The “until I decide” language keeps the door open, but the initiative to return is entirely with him.


Profile

FieldValue
Sage IDJosephB
TitleVP of Product, Gen AI
IndustryInsurance (Fortune 500)
Company Size10,000+
HQWisconsin
Dev population~2,400 developers
AI toolsGitHub Copilot (primary), Databricks Assistant, custom agents, Claude Code (available but less used)
Sagetap tenure5 years (joined Feb 2021), 33 activities, 25+ ratings

Call Timeline

0:00–3:20 — Opening and introductions

JosephB opened by describing his environment: couple dozen AI applications in production, third-party tools, agent automation, standardizing technologies, creating best practices, governance tools, data management tools. He asked what part of the industry Empathic is in.

Bryan: “We build systems of understanding for AI. Initially about helping humans understand what their AI systems are doing, and eventually helping AI understand what humans are doing.”

Paolo gave a longer setup about tools and chaos in engineering organizations.

JosephB at 3:20: “Cool, I think that can break down to a lot of different stuff. So, where do you want to focus the conversation?” — He’s politely saying: I still don’t know what you do. Give me a specific.

3:27–7:13 — Bryan’s monologue (the structural failure)

Bryan delivered a roughly 4-minute continuous pitch covering:

  • The token budget wars thesis (companies blowing 2026 budgets in Q1)
  • Rapid adoption without measurement
  • Engineering coaching and code review decay
  • Git insufficiency for understanding AI-augmented work
  • Technical architecture: structured traces, super structures, unified data from Claude/Codex/Copilot/Cursor
  • Historical data processing and session rehydration

This hit all the right themes in the abstract but didn’t answer JosephB’s question. He asked “where do you focus?” and got a tour of the entire problem space.

7:13–8:11 — JosephB asks for scope

“Are you focusing specifically on the developer environment or broadly across all these cases?” Still scoping. Bryan clarified: primarily AI-driven software engineering.

8:11–13:51 — JosephB gives gold (discovery that should have happened first)

This is the segment where JosephB told the team exactly what he needs.

  • “The number one problem is really figuring out what’s the right metric to measure impact.” He framed a spectrum:
    • Easy case: legacy modernization (6x ROI — invested $100K, saved $600K in time). Clear dollar translation.
    • Hard case: “Should we bring in Claude Code alongside Copilot?” The benefit of one vs. the other is not about direct cost — it’s about marginal value.
  • He described AI over-reliance — production bug where the model fell back to its 2025 training year instead of 2026 on a date form. “These kind of mistakes would never happen before.”
  • Token budget variance: his company has $50/month per developer, while a friend at Datadog has $500/week — two orders of magnitude.
  • The critical ask at ~12:19: “Before we talk to any vendor that presumes to measure it, it’s really about what exactly are you going to measure that makes sense for us?”

Paolo interrupted the discovery at 10:40 with “That’s our wheelhouse” after JosephB described the spectrum from easy-to-measure to hard-to-measure cases. This cut short what was the best discovery moment on the call.

13:51–18:48 — The pitch that didn’t answer his question

Bryan responded to the “what will you measure?” ask with:

  • “You can do some of this yourself locally with Claude Code” (undermines the product)
  • “Attempting to do that at organizational scale is very challenging” (motivation, not answer)
  • “Our day one is not a panacea… it’s that you have a data substrate to begin to ask these questions” (infrastructure positioning)
  • Session rehydration when Anthropic goes down (ergonomic feature, not what he asked)

Paolo added: “No developer friction, it’s a daemon that runs in the background.” (Correct but not what JosephB asked for.)

JosephB’s crystallizing interpretation at 18:48: “So if I understand correctly, you are focused less on measuring the direct ROI and more on helping us understand where we spend money in places where there might be inefficiencies.”

He was generous with that interpretation. It was also a yellow card — he downgraded Pathbase from “product that measures ROI” to “data tool that shows where money goes.”

19:06–22:47 — Vague metrics

Bryan mentioned productive vs. unproductive engineers, token leaderboards being counterproductive, knowledge sharing between effective and struggling engineers. Paolo added prompt-to-PR cycle time.

JosephB pushed back at 22:47: “Very different if you’re debugging a feature that’s already in production versus engineering an entire architecture from scratch.” He’s saying the metrics don’t account for context. Then he asked the scoping question: “How flexible is the platform? You basically say, here’s the data, query it?”

That question — “is this just a data platform to query?” — is the moment JosephB mentally categorized Pathbase as infrastructure, not a product.

23:26–28:14 — The PR review workflow (the best thing said in the call)

Bryan described the PR review flow: developer opens PR, auto-generated link to Pathbase session trace, reviewer can see the full agent decision history, can boot up the agent at the point of a concerning change and interrogate it. “The person reviewing it can interrogate and work with the full body of lineage that went into the creation of work.”

This is the only concrete, opinionated, decision-driving product feature described in the entire call. It arrived at minute 24 of 32.

28:14–32:20 — Honest assessment and soft close

JosephB at 28:22: “My honest take, I think there’s something interesting here. My concern is primarily around the maturity of the product.”

At 28:38: “I think it’s going to require someone in our land to be very proactive and opinionated in figuring out how to use it.”

At 28:47: “We have a VP on our team who’s specifically in charge of AI for the developer experience. I’m going to share this with her.”

At 29:36: “If it’s any encouragement, I think the problem you’re tackling is real.”

At 30:10: “It’s really a question, okay, so what’s the biggest problem right now, and how quickly can we activate a tool to generate that ROI, and how much do we need to invest in customization or building a process around that, or is it something that comes more out of the box?”

Paolo offered a POC. JosephB said he’d follow up after talking to the VP.


What JosephB Revealed

Pain points confirmed

  1. The right-metric problem. He’s not struggling to adopt AI — he’s struggling to measure impact across the full spectrum from easy (legacy modernization, dollar-translatable) to hard (should we add another model family? what’s the marginal value?).

  2. AI over-reliance cost. Production bugs with a different character than pre-AI bugs. The date-falling-back-to-training-year example is a trace-visible failure.

  3. Token budget variance across companies. $50/month vs. $500/week at a peer company (Datadog). He has no framework for knowing whether his allocation is right.

  4. The VP of AI for Developer Experience. Someone on his team owns the developer enablement charter. This is the actual buyer/implementer. JosephB is the strategy leader; she’s the one who would deploy a tool.

Numbers and facts volunteered

  • “Couple of dozen applications already in production” (AI apps)
  • $50/month per developer token budget
  • Peer comparison: Datadog at $500/week per developer
  • Production bug: date falling back to 2025 training year instead of 2026
  • Legacy modernization: 6x ROI ($100K invested, $600K saved in time and manpower)
  • He has a peer VP specifically in charge of AI for developer experience (unnamed)

Signals missed (questions not asked)

  • What company is this? Still don’t know. Wisconsin Fortune 500 insurance narrows it considerably.
  • What do the custom agents do? Never probed despite “custom agents with poor visibility” being highlighted in his qualifying answers as a key signal.
  • How is the domain budget conversation going? The token allocation crisis from his Q1/Q2 answers — the single strongest signal — was mentioned by Bryan in the opening monologue but never explored through JosephB’s experience. Never asked: “When a domain lead pushes back on their token allocation, what happens?”
  • Who is the VP? She was mentioned but never named or characterized beyond “in charge of AI for developer experience.”
  • None of the three universal probes were asked. The mental model reveal, champion chain, and purchase intent probes from the playbook were all skipped — again.

What Worked

1. JosephB is a generous, substantive interlocutor

He gave detailed, structured answers without much prompting. The metric spectrum (easy to hard), the date bug example, the token budget variance, and the honest assessment at the end were all volunteered. Even in a call where the pitch struggled, the signal quality from JosephB was high.

2. The PR review workflow resonated

When Bryan finally described the concrete flow — session trace linked from PR, reviewer interrogates the agent at the point of a concerning change — JosephB engaged. His response (“there’s something interesting here”) came directly after this segment. If this had been deployed at minute 8 instead of minute 24, the call has a different shape.

3. The token budget wars thesis matched his reality

Bryan’s opening about companies blowing their 2026 token budgets landed — JosephB confirmed the dynamic is real at his company. The framing was correct even if the delivery was too long.


What Failed

1. No discovery happened

The first 7 minutes were continuous pitch. JosephB asked “where do you want to focus?” at 3:20 — an open invitation to turn the conversation into discovery — and received a 4-minute monologue instead. The prep note had specific opening questions designed for JosephB (“Can you walk me through how that conversation goes today when a domain lead pushes back on their token budget?”). None were asked.

The playbook’s “weave, don’t stack” directive was ignored. The playbook’s “by minute 15, start transitioning” rule was irrelevant because there was nothing to transition from — it was pitch from the start.

2. JosephB asked “what will you measure?” three times and never got a concrete answer

  • Ask 1 (8:11): “The number one problem is figuring out what’s the right metric to measure impact.”
  • Ask 2 (12:19): “What exactly are you going to measure that makes sense for us?”
  • Ask 3 (20:39): “What actions do they take? What insights did they get, and what did they do with them?”

Each time, the response was architectural (“data substrate,” “super structures,” “heuristics”) rather than concrete (“we measure cost per merged PR, token waste, rework rate, agent efficiency comparison, and review burden”).

3. “Data substrate” positioning killed the product perception

When JosephB asked “is this basically a platform to query?” — that was the moment he categorized Pathbase as infrastructure without opinions. His Sagetap feedback confirms this was the core objection: “just delivers information, but without specific insights, suggested action items, or downstream process.”

Enterprise VPs buy products that answer questions. They don’t buy infrastructure they have to build on top of.

4. The “we’re early” signaling damaged credibility

Phrases like “our early design partners,” “we’re launching with,” “beginning to have a unified data substrate,” and “we can probably get it live for you today” sent mixed signals. JosephB heard: these people are still building this. His feedback — “not mature enough” — may partly reflect the product, but it definitely reflects how the product was described.

5. Two-person pitch without role clarity

Bryan and Paolo both pitched without a defined division of labor. Paolo’s “that’s our wheelhouse” at 10:40 interrupted JosephB’s best discovery moment. His “no developer friction” at 18:15 answered a question JosephB didn’t ask. The pre-call prep didn’t define roles (who discovers, who pitches, who closes). For a joint call, one person should ask questions and the other should provide crisp answers when tagged in.

6. None of the prep was used

The pre-call briefing had specific openers (“You mentioned 2,400 developers and a token allocation problem”), specific probes (“What are those custom agents doing?”, “What does ‘micro’ mean to you?”), all three universal probes, a time-boxed structure, and a listening-signals table. None of it appears to have been referenced during the call. The opening question about the domain budget conversation — designed specifically for JosephB — was never asked.


What JosephB Wanted (post-call synthesis)

Combining his qualifying answers, his on-call questions, and his Sagetap feedback, here’s what JosephB was evaluating for:

  1. Specific, named metrics that Pathbase computes. Not “we capture data” — what questions does the product answer out of the box?

  2. Action that follows from insights. “What did they do with it?” He wanted to hear: “Engineering leads used the rework rate data to identify which teams needed coaching. The VP used the agent comparison data to decide to drop Copilot for Claude Code on backend services.”

  3. Maturity signal. Evidence that the product works in production at scale, not that it’s an interesting research prototype. The POC offer came too late and was framed as exploration, not demonstration.

  4. Low implementation burden. “It’s going to require someone to be very proactive and opinionated in figuring out how to use it.” He didn’t want that. He wanted a product that has opinions built in, not a platform that requires his team to build the “so what” layer.


Comparison to Previous Calls

DimensionJosephBBrett_1949Kevin_5542Joseph_1446
Discovery qualityLow — almost none, pitch-firstMedium — constrained by outbound formatHigh — Kevin volunteered extensivelyLow — pitch came first
Pitch qualityWeak — abstract, infrastructure-positionedMedium — Minority Report concept landedMedium — compressed but concepts landedWeak — monologue, missed probes
Prospect signal qualityHighest — specific metrics, real numbers, concrete examplesHigh — null-check bug, production painHigh — code review gap, 6 initiativesHigh — vanity metrics, Datadog model
OutcomePassed (soft no)Demo follow-up securedSoft close, “I’m intrigued”Mixed — strong signal, weak execution
Key lessonLead with concrete metrics, not architectureGive him space, don’t monologueWeave product into discovery, close specificallyDon’t pitch first, do probe

Pattern across all four calls: The prospect signal quality is consistently high. The pitch quality is consistently the bottleneck. Every call produces rich discovery signal despite the execution problems, which means the campaign is reaching the right people. The conversion failure is in the pitch, not the targeting.


Action Items

Immediate

  1. Build the pitch book — Paolo proposed this in the debrief and he’s right. Five specific questions Pathbase answers, with examples. Rehearsed one-liner, one-paragraph, five-minute version.

  2. Prepare a follow-up for JosephB if a channel exists. If Sagetap allows post-Pass messaging or if he circles back: “We heard your feedback. Here are five specific questions Pathbase answers for engineering leaders. Would it be worth 20 minutes to show [VP name] what a proof of concept looks like?”

  3. Send a video demo if possible. Bryan offered to prepare one. JosephB didn’t explicitly accept or decline. If Sagetap allows it, send.

Before next call

  1. Define roles for joint calls. If Paolo joins again: one person asks questions (discovery), one person provides crisp product answers when tagged. Don’t dual-pitch.

  2. Print the five metrics. Before every call, have the concrete answer to “what will you measure?” on a card. Cost per merged PR. Token waste. Agent comparison. Rework rate. Review burden. Say them by name when asked.

  3. Use the prep notes. The JosephB prep had an excellent opening question, specific probes, universal probes, and a listening-signals table. None were used. Treat the prep as a literal script for the first 5 minutes, then freestyle once the prospect is talking.

Strategic

  1. The “data substrate” framing is retired. Every subsequent pitch leads with decisions the product enables, not infrastructure it provides. “We answer these questions” not “we give you a platform to query.”

  2. JosephB’s date bug is the new best call story. More concrete than Brett’s null-check bug because it’s a novel failure mode (model training year leaking into production data) that only AI-generated code can produce, and it’s trace-visible.