Sagetap CTO Platform Campaign — Inbound Evaluation (First Batch)
Campaign: CTOs: Agent Sessions Are Becoming Your Company’s Institutional Memory — Are You Using Them? Campaign type: Market Awareness (30 min) Date evaluated: 2026-06-19 Applications received: 2026-06-15 Campaign spec: 2026-06-12-sagetap-cto-platform-campaign Format reference: 2026-05-14-sagetap-inbound-evaluation
The headline
The campaign that says CTOs drew two CISOs and zero CTOs in its first batch, and both arrived with active AI-security initiatives — runtime guardrails, prompt-injection defense, DLP, shadow-AI discovery — that Pathbase does not build. This is the Archetype 2 governance/security magnet that has now appeared in every Empathic Sagetap campaign: 5 of 11 in the original batches, and 2 of 2 here, under a title aimed explicitly at the technical founder and a register (institutional memory, the record of why) deliberately built to repel the security frame. It did not repel it.
Two things are true at once, and the second is the one worth sitting with:
- On targeting, the title is mis-firing. “Are You Using Them?” plus “record” plus “audit-adjacent” language reads to a CISO as agent oversight / monitoring / forensics — their job — not as the data foundation of an AI-native company. If actual CTOs are the goal, the title is selecting against them.
- On thesis, the wrong buyer validated us anyway. Both CISOs, unprompted, described the record-of-why almost verbatim from the campaign’s own framing — Link asked “why was this change made, what alternatives were rejected,” Leopold wanted to “examine the issues agents are encountering, identify themes… make the agents more reliable.” The institutional-memory thesis is legible to security buyers under a forensics/agent-ops framing. That is real signal, and it is exactly what a free data instrument is supposed to surface.
Neither call clears the accept bar (reasoning below). Both are pure data, banked at zero credit cost. The campaign is working as designed; the design just surfaced something about who the framing attracts that needs a decision before the next batch.
Summary
| # | Sage | Role | Industry | Size | HQ | Call cost | Fit (this campaign) | Archetype | Key signal |
|---|---|---|---|---|---|---|---|---|---|
| 1 | Link | CISO | Insurance | 10,000+ | France | 1,500 | 🟠 DECLINE call / 🟢 HIGH data | Archetype 2 (security forensics) | Strong evaluator, thesis-aligned answers, but initiative is AI-app security — wrong product |
| 2 | Leopold_R | CISO | Manufacturing | 1,001–5,000 | Texas, US | 850 | 🔴 DECLINE call / 🟡 MOD data | Archetype 2 (AI governance) | Machine-consumer Q2 answer, but ~3 agent users and a pure DLP/shadow-AI initiative |
Fit reading for this campaign is two-dimensional. Calls cost credits and the default is decline, so the call column is mostly red. The data column is what the campaign exists for, and it is where the value sits.
Qualifying Questions (for reference)
The CTO Platform campaign’s three questions (2026-06-12-sagetap-cto-platform-campaign):
- Q1 — Retention & felt loss: Roughly how many engineers work with AI coding agents daily, and what happens to those sessions after the work ships? Are they retained anywhere, and if you needed to go back to one to understand why a change was made, could you?
- Q2 — First query & consumers: Imagine a complete, searchable record of every agent session — intent, alternatives, dead ends — joined to PRs, deploys, incidents. What is the first question you would ask it, and who or what else would use that record?
- Q3 — Commitment conditions: What would have to be true for agent session history to become a first-class system of record? Where would the data live, what would need to be redacted, and would you build or buy?
Individual Evaluations
Link
CISO, Insurance, 10,000+, France — 🟠 DECLINE the call (1,500 credits), 🟢 HIGH-value data
| Field | Value |
|---|---|
| Seniority | CXO (CISO) |
| Industry | Insurance |
| Company size | 10,000+ |
| HQ | France |
| Sagetap rating | 4.9 (109 ratings) |
| Badges | Verified, Heavy Evaluator |
| Activity | 167 activities, 4 initiatives, joined Aug 2024 |
| Tenure in role | 4+ years (Jan 2022–now) |
| Role | CISO. Evaluates emerging tech, aligns with strategy, oversees pilots, assesses risk across 22+ entities. PCI, GRC, PMQ certified. |
Qualifying Responses (verbatim)
Q1 — Retention & felt loss:
I would not claim one/a single exact number across 27 entities. The usage is growing, but our maturity is mixed and not always centrally visible. In many cases the session history stays in the tool or disappears from the engineering record. That is the gap: after code ships, we often have PRs, tickets, and commits, but not the full agent reasoning or decision trail.
Q2 — First query & consumers:
My first question would be: why was this change made, what alternatives were rejected, and did the agent touch sensitive logic or data paths? Security, engineering leads, reviewers, incident responders, auditors, and future agents could all use it. The value for us is faster review, better onboarding, stronger incident analysis, and more evidence when something breaks.
Q3 — Commitment conditions:
For us it would need to sit in a controlled enterprise environment, ideally our cloud or a tenant we can govern. Capture by default would require clear controls for source code, secrets, PII, customer data, retention, access rights, and audit trail. I would expect a buy plus integration model, not a full in house build, unless the data risk was too high.
Assessment
These are the most thesis-aligned qualifying answers the portfolio has produced since Joseph_1446, and they came from a security buyer. Read Q1 again: “we often have PRs, tickets, and commits, but not the full agent reasoning or decision trail.” That is the vision line — “Git keeps what shipped and throws away how it was made” — restated by a stranger in his own enterprise vocabulary. Q2’s first clause (“why was this change made, what alternatives were rejected”) is the product pitch nearly verbatim. Q3 is a sophisticated, buy-oriented, BYOI-shaped answer: governed tenant, explicit redaction categories, buy-plus-integration. If you scored these three answers blind, you would call this a strong-fit platform buyer.
Then you read the initiative, and the fit collapses. His active initiative is “AI Application Security & Agent Security” — securing AI apps “at build time, test time, and runtime,” targeting data leakage, prompt injection, unsafe outputs, policy violations, excessive permissions, runtime guardrails, red teaming. He is shopping for a runtime AI-security platform. Pathbase does none of that. He listed Pathbase among 20 “Interested” solutions — alongside Bloom Security, AIFT-Vulcan, NetFoundry, Revenium — because it mentions agents and an audit trail, and his net is wide. The third clause of his Q2 (“did the agent touch sensitive logic or data paths”) is the tell: he is reading the session record as a security forensics artifact, not as engineering institutional memory.
Why the call is a decline despite the 4.9 rating and 10,000+ scale. Taking this call (1,500 credits — the most expensive in either batch) is the JosephB failure pre-loaded. He arrives expecting runtime protection, red teaming, and policy enforcement; we show a session-trace record; he rates us “interesting but not what I need,” and we have spent the batch’s largest credit block to learn what the free answers already told us. The accept bar exists for exactly this case: a named, urgent, well-funded initiative whose category is not ours. Criterion 4 (named initiative + stakeholder chain matching what Pathbase ships) fails on the matching clause.
Why the data is nonetheless the most valuable in the batch. Link is the single best evidence yet that the record Pathbase captures is the substrate a security buyer wants for AI forensics — even though Pathbase doesn’t do the detection layer on top. “Why was this change made, what alternatives were rejected, did it touch sensitive data paths” is a forensics query against precisely the DAG Toolpath stores. He is a Heavy Evaluator at a 22-entity insurer with a Q2-2026 budget and urgency. He is the wrong buyer for the product as it ships and a near-perfect buyer for a security/audit projection of it that does not exist. That tension is the strategic finding (see synthesis).
Initiative — “AI Application Security & Agent Security” (AI Security & Governance)
- Phase: Evaluating | Type: New Purchase | Target: Aug 2026 | Urgency: “expect this area to move in Q2”
- Scope: reduce practical security risk of AI adoption across 60+ AI/ML projects in 22+ entities. Explicitly not cloud posture or general GRC — securing AI-driven apps at build/test/runtime.
- Requirements: data leakage, prompt injection, insecure tool use, unsafe outputs, policy violations; preventive + detective controls; runtime guardrails, monitoring, red teaming; MS-led environment, multi-entity rollout.
- Considering 20 solutions (NetFoundry, Tessl, Pathbase, Browserbase, Shelf, Bloom Security, Willow, AIFT-Vulcan, Revenium, +others). 2 meetings completed.
- Pathbase relevance: LOW for the initiative as scoped. The initiative is a runtime AI-security PoV. Pathbase is a capture-and-record platform. The overlap is the word “agent” and the concept “audit trail,” not the buying need. He’d evaluate Pathbase against runtime-protection criteria it can’t meet.
Recommendation
Decline the call. Bank the answers as the batch’s highest-value data. If a zero-cost channel opens (Sagetap message, or he resurfaces), the honest framing is narrow and true: “We’re not a runtime AI-security control — we don’t do prompt-injection defense or red teaming. What we keep is the decision-and-reasoning record your Q1 said you’re missing, in a governed tenant with the redaction controls you listed. If your incident responders or auditors ever need the why behind an agent change, that’s us.” That sorts instantly whether there’s a forensics-substrate conversation or not, at no credit cost. Do not spend 1,500 credits to discover he wanted Bloom Security.
Leopold_R
CISO, Manufacturing, 1,001–5,000, Texas — 🔴 DECLINE the call (850 credits), 🟡 MODERATE-value data
| Field | Value |
|---|---|
| Seniority | CXO (CISO) |
| Industry | Manufacturing |
| Company size | 1,001–5,000 |
| HQ | Texas, US |
| Sagetap rating | 4.2 (38 ratings) |
| Badges | Verified, Evaluator |
| Activity | 103 activities, 1 initiative, joined Jan 2026 (newer) |
| Tenure in role | 7 years (May 2019–now) |
| Role | CISO, fully accountable for IT Security & Risk. M&A-active company. Sole decision-maker / purchaser for cybersecurity tech. Program embracing zero trust, AI mitigations, cloud security, M&A. |
Qualifying Responses (verbatim)
Q1 — Retention & felt loss:
~3. There is nothing preventing deletion of work sessions, so there is not a written record to refer back to later.
Q2 — First query & consumers:
Interesting question - we would likely look for ways to improve the agents, so we’d start by examining the issues agents are encountering and identify themes so we can address them. The goal would be to make the agents more reliable, so they require less oversight. Probably all of the other roles you mentioned would use it although auditors may not use it at the moment. Moreso engineers and reviewers.
Q3 — Commitment conditions:
I need to know what agent is, what tasks are assigned to each, what each agent has done and at what time and be able to tie that to its task(s). This would aid us in identifying what has gone wrong with an agent and why so the appropriate adjustments can be made. In terms of where that data can be hosted and whether we build/buy, we’re open.
Assessment
His Q2 is the most platform-register answer in the batch — and it comes attached to the smallest engineering footprint. “We would likely look for ways to improve the agents… make the agents more reliable, so they require less oversight” is a machine-consumer answer in the campaign’s exact sense: the record’s first job is to make the agents better, not to police people. That is the vision second-register belief (agents learning from the organization’s own experience) surfacing unprompted, and it’s the answer the Q2 probe was built to detect. The campaign note flagged the machine-consumer percentage as “the single most important number this campaign produces” — Leopold is the batch’s one tick in that column.
But the scale and the initiative gut the call case. Q1 is “~3” engineers using agents. At three users there is no corpus, no aggregate-analytics value, and the “improve the agents” use case is a weekend of reading three people’s sessions by hand, not a platform purchase. Accept-bar criterion 1 (machine-consumer answer plus 50+ daily agent users) fails hard on the second clause. And his only initiative — “Comprehensive AI Security, Governance, and Risk Management” — is pure Archetype 2: GenAI governance, prompt defense, DLP, shadow-AI discovery (rogue MCP servers), NIST CSF 2.0, scoped to “all knowledge workers.” His Q3 wants agent inventory: “what agent is, what tasks are assigned, what each has done and at what time.” That is an agent-observability/governance dashboard for a security org, not an engineering system of record.
The read. Leopold is a newer evaluator (4.2, joined Jan 2026, single initiative) at an M&A-driven manufacturer, looking for an AI-security-and-governance suite, who answered the Q2 thought experiment with genuine curiosity and landed on the agent-improvement framing because that’s the honest answer for someone with three agent users and no detection product yet. The answer is encouraging for the thesis and useless for a call: tiny scale, wrong initiative category, no purchase path for what Pathbase ships.
Initiative — “Comprehensive AI Security, Governance, and Risk Management” (AI Security & Governance)
- Phase: Evaluating | Type: New Purchase | Target: Aug 2026
- Scope: secure the GenAI roadmap — governance, prompt defense, data privacy, NIST CSF 2.0. LLM usage + agentic AI the primary concern.
- Requirements: enterprise-wide visibility into AI use; monitor employee LLM prompts for DLP; shadow-AI discovery (rogue MCP servers); platform-agnostic agent security; agent accountability (who owns which agent, tasks, permissions); ROI dashboarding.
- Considering 15 solutions (Gleam AI, Arambh Labs, Browserbase, +others). 1 meeting completed (5 total meetings).
- Pathbase relevance: NONE for the initiative as scoped. DLP, shadow-AI discovery, and prompt monitoring are a different product. The only thread to Pathbase is the agent-accountability requirement (who did what, when, tied to tasks) — which Toolpath captures structurally — but he wants it as a security control, not an engineering record.
Recommendation
Decline the call. Smaller, newer, ~3 agent users, pure security-governance initiative, no scale and no matching purchase path. Banking the Q2 answer is the entire value: it is the machine-consumer data point, and it confirms the thesis is legible even at the bottom of the market. Spend zero credits.
Cross-Cutting Synthesis
1. The Archetype 2 magnet is now a cross-campaign constant, not a campaign-specific miss
This is the finding that outranks the individual evaluations. Empathic has now run three Sagetap campaigns, and the governance/security buyer has dominated or co-led every one:
| Campaign | Register / title | Archetype 2 share of inbound |
|---|---|---|
| Original “See What Your AI Agents Are Doing” | Visibility | 5 of 11 (45%) across first batches |
| CTO Platform “Institutional Memory / Are You Using Them?” | Platform / data foundation | 2 of 2 (100%) first batch |
The platform campaign was designed to repel the security frame — the persona text explicitly excludes “surveillance, DLP, or agent inventory” buyers, and the register is institutional memory, not oversight. It pulled two CISOs anyway, both with DLP/runtime-security initiatives. At n=2 this is not yet a distribution, but the direction is unambiguous and consistent with everything prior: the strongest, most repeatable inbound demand Empathic generates is from security/governance buyers for a product it has chosen not to build. The 2026-05-28-sagetap-campaign-revision rule — “don’t build a governance campaign until you ship governance features” — was correct as written, but the evidence for revisiting it is now mounting from a third independent direction. This deserves an explicit strategic decision, not another round of declining and moving on. (Filing as a flagged follow-up, not acting on it here — it’s a product-strategy call, not an evaluation output.)
2. The thesis is legible to the wrong buyer — which is the most useful thing this batch proved
Both CISOs described the record-of-why in their own words before being pitched. Link: “the full agent reasoning or decision trail… why was this change made, what alternatives were rejected.” Leopold: “examine the issues agents are encountering, identify themes… make the agents more reliable.” The category claim in the listing (“the data foundation of an AI-native company”) landed — just with a buyer who reads it through a security/ops lens and routes the value to forensics and agent-reliability rather than engineering memory. The language-resonance row of the campaign’s data plan comes back positive: “record,” “reasoning,” “decision trail,” “improve the agents” were all adopted readily. The thesis is communicable. The open question the campaign was built to answer — does the register pull the right buyer — comes back not yet, on this evidence.
3. The title verb probably amplified the security pull — but the live campaign is locked
The live title is ”…Are You Using Them?” — different from the recommended ”…Are You Keeping Them?” “Keeping” is a retention/memory verb; “using” is an operational/oversight verb, and to a CISO “using agent sessions” reads as monitoring, governance, and forensics — exactly the frame both applicants brought. The hypothesis is worth testing.
But it can’t be tested by editing this campaign: Sagetap campaigns are locked once live. The current campaign will keep running as-is, with this title, this Market Awareness type (note: also different from the Research Calls the spec recommended — that divergence is now frozen too), and these questions. Testing the verb therefore means standing up a separate campaign with the “Keeping” framing and running it alongside this one. Because campaigns and inbound are free, the test costs nothing in credits — the only cost is another listing to monitor and the mild portfolio fragmentation of a fifth concurrent campaign. The lever is real; the mechanism is “launch a sibling,” not “fix a word.”
4. Hosting and build-vs-buy data (the rows that transfer regardless of archetype)
Even from the wrong buyers, the Q3 commitment data is real and on-thesis for the BYOI architecture:
- Hosting: Both want governed/flexible hosting. Link: “our cloud or a tenant we can govern.” Leopold: “open.” Zero demand for pure vendor-hosted-only; the governed-tenant / BYOI posture the architecture already assumes is confirmed, n=2.
- Redaction: Link enumerated the categories unprompted — source code, secrets, PII, customer data, retention, access rights, audit trail. This is a clean, free prioritization input for the redaction roadmap (already the top blocker for auto-upload).
- Build vs. buy: Link “buy plus integration, not in-house build”; Leopold “open.” No build-threat signal from either — which, for security buyers at this scale, is mildly notable (they’re not planning to roll their own session record).
Aggregate Data Tally (first batch, n=2)
Updating the campaign’s collection plan with this batch. Two data points is not a distribution — recording to accumulate, not to conclude.
| Data point | Batch reading (n=2) |
|---|---|
| Daily agent users | 27 entities / “mixed, not centrally visible” (Link); ~3 (Leopold). One enterprise-fragmented, one trivial. |
| Retain sessions today? | Neither. Link: “stays in the tool or disappears.” Leopold: “nothing preventing deletion.” 0/2 retain. |
| Homegrown retention tooling | 0/2. |
| Retrieval-failure anecdote | Link implicitly (can’t reconstruct the why from PRs/tickets/commits). 1/2 felt, 0/2 concrete story. |
| First-query corpus | ”Why was this change made / what was rejected / did it touch sensitive paths” (Link, forensics); “what issues are agents hitting, by theme, to improve them” (Leopold, agent-reliability). |
| Human vs. machine consumer | Leopold = machine-consumer (improve the agents). Link = human-consumer (security/reviewers/auditors) + “future agents” as a list item. 1 clear machine-consumer of 2. |
| Hosting requirement | Governed tenant / own cloud (Link); open (Leopold). 0/2 require pure on-prem; 0/2 accept vendor-only-no-governance. |
| Top redaction categories | Source code, secrets, PII, customer data, retention, access rights, audit trail (Link). |
| Build-in-house inclination | 0/2. Link explicitly buy; Leopold open. |
| Language resonance | Positive. Both adopted record / reasoning / decision-trail / improve-the-agents readily. |
| Archetype mix | 2/2 Archetype 2 (security/governance). 0 CTOs, 0 platform/data-foundation buyers, 2 CISOs. |
Recommendation
Decline both calls. Bank both as data. Total credits spent: 0 of 2,350 available. That is the campaign performing exactly to its design — the answers are the product, and both arrived free.
The live campaign rides as-is. Sagetap campaigns can’t be edited once live, so there is no tuning to do here — the title, the Market Awareness type, and the questions are frozen. That’s fine: it’s free, it’s producing data, and even Archetype 2 data is useful accumulation. Let it keep running and keep harvesting answers.
The only lever is a new campaign, and that’s a decision, not a fix. If the verb/framing hypothesis is worth testing, the mechanism is a separate sibling listing — “Keeping” instead of “Using,” and a title led by the engineering consequence (“the why behind your code”) rather than the asset noun (“institutional memory,” which a security buyer hears as their mandate). It’s free in credits. The real question is whether a fifth concurrent campaign is worth the attention, or whether the cleaner move is to wait, let this one accumulate a real n, and fold the framing lesson into the next campaign you’d build anyway. My lean: don’t launch a sibling just to test one word on n=2 — let this ride, and carry the “Keeping / engineering-consequence-forward” framing into the next net-new campaign. Revisit if this one keeps coming back pure Archetype 2 across a real batch.
What would change the call decision in a future application: an actual CTO / technical founder / chief architect (not a CISO) whose Q2 names their own systems and routes the record to engineering memory or agent learning, at 50+ daily agent users — i.e., the accept bar as written. Neither applicant here is close, and the gap is not marginal: both are the wrong role, with the wrong initiative category, and one is at trivial scale.
The strategic item that outlives this batch: the Archetype 2 pull has now recurred across all three campaigns, including one engineered to deflect it. That is no longer a targeting nuisance to route around — it is a consistent market signal that the most reliable inbound demand Empathic creates is for agent security/governance/forensics. The tactical answer is unchanged (don’t build a security product on n=2, don’t burn calls on mismatched initiatives). The accumulation is the point, and it belongs in front of the team as a deliberate “are we sure we’re not building this?” decision rather than as an evaluation footnote.
Related
- 2026-06-12-sagetap-cto-platform-campaign — Campaign spec, accept bar, and data-collection plan this batch feeds
- 2026-05-14-sagetap-inbound-evaluation — First-batch format reference and origin of the two-archetype framing
- 2026-05-28-sagetap-campaign-revision — The “don’t build a governance campaign yet” rule now under accumulating pressure
- 2026-05-29-josephb-sagetap-call — The maturity/wrong-expectation failure the Link decline is built to avoid
- 2026-06-12-single-timeline-thesis-pathbase-alignment — The platform register both CISOs validated in their own words
- vision — “Git keeps what shipped and throws away how it was made,” restated by Link’s Q1
- Pathbase — Product context for the security-forensics-substrate tension