Kevin_5542 (CTO/CISO) — Sagetap Call & Debrief
Campaign: Engineering Leaders: See What Your AI Coding Agents Are Actually Doing Platform: Sagetap ($1,000/call — inbound) Date: 2026-05-20 Outcome: Kevin will review and follow up. Intrigued but wants to evaluate further. No demo explicitly scheduled.
Profile
| Field | Value |
|---|---|
| Name | Kevin (surname TBD — LinkedIn not exchanged on call, he said he’d send details after) |
| Sage ID | Kevin_5542 |
| Title | CTO / CISO |
| Company | AI enterprise software company (unnamed on call) — serving Retail, Fintech, Energy, Gov, Healthcare |
| Company size | 51–200 employees |
| Location | Dublin, Ireland (at time of call; profile says UK) |
| Sagetap | Joined several years ago, spiky activity, 23 meetings completed, 4.8 rating, Heavy Evaluator |
| Initiatives | 6 active (see 2026-05-14-sagetap-inbound-evaluation#Kevin_5542) |
What his company does
Professional services / co-creation firm building AI enterprise software for large clients (1000+ person orgs). They do client-embedded development — working inside client codebases, co-creating features, building agentic workflows. This means they care about code quality and client-facing standards (naming conventions, structure, compliance with client coding standards).
Key context from qualifying responses (pre-call)
- Claude Code is ~80-85% of their AI agent usage. Copilot used when co-creating with Microsoft-ecosystem clients. Some Codex.
- Production deployments, not just R&D — 50% on production work, rest on R&D/prototypes
- “None have solved the problem for us yet” on AI-powered PR review tools
- “Agentic AI platform gaps” initiative targeting Jun 2026 — next month
- 6 active Sagetap initiatives, assembling a full AI infrastructure stack
Green/Red Flag Scorecard (post-call)
Green flags
- He manages a team of software engineers who write code — professional services firm co-creating with clients, engineers in agentic mode daily
- AI coding agents are adopted — Claude Code is ~80-85%, Copilot for client co-creation, some Codex. Production work, not just experiments
- He articulates the visibility or coaching gap in his own words — described code review bottleneck, agents missing duplicate functions, “trust” as the core challenge
- He engages with the product concept — strongly engaged with dead-ends/organizational knowledge concept, Jellyfish comparison, “two doors” automated PR routing vision
- He asks how Pathbase works — not specifically; he received the pitch late in the call
- He names someone else who’d care — described Digital Performance Team and designers as adjacent users, though didn’t name a specific person
- Has a budget or initiative forming — 6 active Sagetap initiatives, Jun 2026 target on most relevant one. Budget exists but not validated on call
Red flags
- None triggered. Genuine engagement throughout. The only concern is the soft close — “I’m intrigued, I’ll review and follow up” is weaker than Brett’s “let me talk to my senior engineers.”
Qualifying Responses (from live conversation)
Qualifying questions were submitted in advance (inbound Sage). The call confirmed and deepened:
Q1 — AI agent adoption (confirmed and expanded): Claude Code is the dominant tool — “80, 85% of the work is going through Claude Code.” Copilot used in co-creation with clients (“it’s a big Microsoft worldview… even though it’s just not as strong and can be frustrating for all sorts of other reasons”). A few people using Codex but “hasn’t really taken off.” This is a Claude Code-primary shop. Kevin noted Opus 4/5 as the inflection point when things changed dramatically.
Q2 — Visibility and ROI (confirmed): Not directly probed on call. His pre-call answer (“qual data showing tremendous benefits in both HOW we work and in WHAT we work on”) went unexamined. This was a missed opportunity — the “how we work” framing is almost verbatim Pathbase’s value prop.
Q3 — Code review and coaching (massively expanded): This was the meat of the call. Kevin described multiple layered problems:
- Volume + size problem: PRs are coming through faster and bigger (“potentially hundreds of commits”), creating a bottleneck where senior reviewers either slow everything down or skip reviews (“oh yeah, this, it’s probably okay”)
- AI code review tools failing: They’ve tested them rigorously — deliberately inserting bad code to see if tools catch it. Results: tools miss a lot, are “overly critical” or not critical enough depending on model updates, lose credibility with engineers who end up manually reviewing anyway
- Agents miss cross-codebase context: Engineer A in full agentic mode rewrites a function that already exists elsewhere, duplicates UI logic, misses business rule overlap in other parts of the codebase. “The agent didn’t pick it up… then it gets missed during the code review cycle”
- Rules vs. model updates: They build guardrails and parameters around agent capabilities, but when the model updates, the rules become wrong — “too restrictive for the new models”
- Trust erosion: “Trust is the word I use… you end up manually reviewing anyway” — echoes his pre-call answer about traditional devs building trust
New signals from the call:
The Digital Performance Team: Non-engineer team using code for business analytics — Google Analytics feeds, product signals, checkout funnel analysis, A/B tests, user frustration detection. They’re “using code in a different way” with MCP and third-party tools, and Kevin wants engineering rigor applied to their work too. They’re currently “just hacking things to death and hoping it works.” Designers are also writing first versions of features in code.
The “Two Doors” vision: Kevin’s articulated future state — a system that can assess whether a PR meets a certain quality bar and either auto-approve (door A: non-contentious, well-executed, agent did it right) or route to human review (door B: needs scrutiny). He described this unprompted as what he’s looking for in tooling. This is the automated quality triage that Pathbase’s annotation layer could eventually enable.
The Jellyfish comparison: Kevin brought up engineering intelligence tools (specifically Jellyfish and similar) as adjacent-but-disappointing. His complaint: “you end up having to tag and procure your code base in such a way that becomes a very rigid thing.” He connected with Bryan’s pitch specifically because it avoids that rigid tagging approach — learning from the process rather than requiring upfront configuration.
The breadth argument: Kevin’s closing insight was powerful — he wants engineers to have breadth across the entire codebase so they can spot where agents go wrong. “The agents are going to be good at reading all the code, but I also want my engineers to be really good at looking at the breadth of the code base.” The trace data from Pathbase could be the mechanism for developing that breadth: seeing how other engineers (via their agent sessions) work across the codebase.
Debrief
Overall Assessment: 🟢 GOOD CALL — worth the $1,000
Fit: 🟢 STRONG. Kevin is the closest thing to ICP in the entire Sagetap batch. CTO of a ~100-person AI software company, Claude Code as primary tool, production work at 50%+, co-creating with enterprise clients. He’s the decision-maker, the budget-holder, and the technical evaluator all in one. His problems map directly to Pathbase’s value prop — code review failure, cross-codebase context loss, trust erosion, rigid tooling alternatives.
Purchase timeline: NEAR-TERM. “Agentic AI platform gaps” initiative targeting Jun 2026 — next month. He has 6 active purchasing initiatives. He’s a serious buyer who uses Sagetap to source tools (23 meetings completed, 4.8 rating). Unlike Brett (pre-budget, 6-9 months), Kevin is actively spending and evaluating right now.
Next step quality: SOFT. “I’m intrigued… we’ll see about taking it forward” is genuine but non-specific. No demo scheduled, no internal stakeholder introduced, no concrete action item. Compare to Brett’s “let me talk to my senior engineers” — that was a specific action with a specific audience. Kevin’s close is vaguer. The follow-up message needs to create a concrete reason to re-engage.
Key Signals
1. The Code Review Failure Loop (strongest signal)
Kevin described a complete trust collapse cycle with AI code review tools:
- Team needs AI code review because PR volume and size have exploded
- They evaluate tools, deliberately inject bad code to test them
- Tools miss things, or catch things inconsistently
- Engineers lose trust: “it means now what I’m doing is I’m getting the tool to review something, now I have to review the output of the tool — why am I doing that?”
- Team reverts to manual review
- Manual review can’t keep up with volume → either bottleneck or skipping
- Back to step 1
Why this matters for Pathbase: Pathbase isn’t an AI code review tool — it’s the substrate that makes AI code review actually work. The reason current tools fail is they review the output (the diff) without understanding the process (the agent’s decision chain). Kevin’s tests — deliberately inserting bad code — are testing whether the review tool can reason about intent, not just pattern-match on code. Traces give the review tool (or the human reviewer) the context to actually catch things.
2. The Cross-Codebase Blindspot
“Engineer A is working on something over here in full agentic mode… the bots will rewrite a function that’s already there somewhere, or they’ll duplicate something, or they’ll miss if we’ve got some part of the UI where we’re upgrading a workflow but that similar functionality is elsewhere in the code base.”
This is a different pain than Brett’s null-check bug. Brett’s was about an agent making a wrong decision within a function. Kevin’s is about agents being unaware of the rest of the codebase — creating duplication and inconsistency that neither the agent, the engineer, nor the code review tool catches until QA. Cross-session traces across an engineering team would reveal these patterns: “Agent A and Agent B both generated implementations of the same business logic in different parts of the codebase.”
3. The Rules Fragility Problem
“We build rules and parameters around the agents… but then the model starts getting updated. Now all our rules are wrong… too restrictive for the new models.”
This is a novel signal not heard in other discovery calls. Static guardrails break when models change. Kevin’s team is sophisticated enough to be building agent-specific rules — and sophisticated enough to see those rules decay with every model update. Trace-based understanding is model-agnostic; it captures what actually happened regardless of which model version produced it. This is an argument for Pathbase that resonates with teams deep enough into agentic coding to have already built (and broken) their own guardrails.
4. The “Two Doors” Automated Triage Vision
Kevin described the system he wants to build: PRs that meet a quality bar go through fully automated (no human review), while PRs that don’t go to human review. He framed this as the future of code review — not better AI reviews, but an AI-powered triage layer that knows when human review is and isn’t needed.
Why this matters: This is exactly what Pathbase’s annotation layer could enable. The “12 features that predict whether this session produced good work or expensive babysitting” from the vision doc — that’s Kevin’s quality bar for door A vs. door B. He described the product roadmap without knowing it exists.
5. Non-Engineer Adoption Signal
The Digital Performance Team and designers writing code is an expansion vector. Kevin wants the same engineering rigor applied to their work but acknowledges they’re “not deep coders.” Trace visibility for non-traditional engineers is a use case that could differentiate Pathbase from dev-only tools.
Honest Feedback on Call Execution
What worked:
-
Let Kevin talk. Kevin is a talker — detailed, tangential, generous with context. Bryan gave him space. The best signals (cross-codebase blindspot, two doors, rules fragility, digital performance team) all came from Kevin having room to riff. This is a major improvement from the Brett call where Bryan’s monologues ran 2-3 minutes.
-
The accordion problem framing. Kevin connected with it. It named something he was experiencing but hadn’t articulated: the expanding/compressing information problem of using one AI to generate and another to review.
-
Peer-to-peer tone was right. CTO to founder, not vendor to buyer. Mentioning the co-founder’s Google background was a good credibility move without overselling.
-
The dead-ends concept landed. Kevin’s response — about organizational knowledge being lost when engineers go down paths and come back — was one of the most engaged moments. He restated it better than Bryan pitched it: “there’s great knowledge in that, in just learning the code base and getting people up to speed.”
What to improve:
-
The pitch came too late and too compressed. The first 22 minutes were pure discovery (good), but the actual Pathbase pitch didn’t start until ~23:00 in a 30-minute call. Bryan had about 4 minutes to explain Toolpath, Pathbase, the graph concept, the cadaver/Minority Report feature, and the business model. It was rushed. Kevin didn’t have time to process or ask questions about the product. Compare to Brett’s call where the pitch started at ~26:00 but the call ran longer. Here, the call hit its hard stop and Kevin left with a concept but not a clear picture of the product.
Fix: Set a mental timer. By minute 15, start transitioning. “I want to make sure we have time to show you what we’re building — can I give you the 5-minute version?” This gives 10+ minutes for pitch, questions, and close.
-
Didn’t probe the two strongest pre-call signals. Kevin’s pre-call answers included two of the sharpest phrases in the batch:
- “None have solved the problem for us yet” — Bryan never probed what they tried and why it failed
- “We’re less focused on quant data because our qual data is showing tremendous benefits in both HOW we work and WHAT we work on” — Bryan never probed what “how we work” means to Kevin
Both are natural entry points to Pathbase’s value prop. “You said none have solved code review for you — what have you tried, and where specifically did they fall short?” would have set up the pitch perfectly.
-
The universal probes weren’t asked.
- “If I showed you a structured trace of everything one of your AI agents did yesterday — what would you look for?” — Not asked. This would have revealed Kevin’s mental model (Archetype 1 vs. 2) with precision.
- “Who else in your org would care about this data?” — Not asked directly. Kevin mentioned the Digital Performance Team organically, but the explicit question would have surfaced the internal champion chain.
- “What would you need to see to put budget behind this?” — Not asked. This is the most important question for a buyer with 6 active purchasing initiatives and a Jun 2026 target. Kevin is spending right now. Asking this would have clarified whether Pathbase is in his evaluation set or a future consideration.
-
“Pathways” instead of “Pathbase” again. Near the end: “pushing that up to pathways.” Same slip as the Brett call. If Kevin searches for “Pathways” he won’t find anything.
-
“Any day now” on launch timing. Kevin asked “months or weeks?” and Bryan said “any day now… if you’re like, I need this today, we can probably get it live for you today.” This is a strong commitment. If Kevin follows up in a week, the product needs to be ready. Expectation management risk.
-
No close on LinkedIn. Bryan asked to connect on LinkedIn. Kevin deflected with “I’ll send you the details after.” That’s a polite no-for-now. The follow-up message on Sagetap needs to include a LinkedIn request or a specific demo invite — something that re-engages before Kevin’s attention moves elsewhere.
Where Kevin Fits in the Sagetap Batch
| Dimension | Kevin (AI Software, Ireland) | Brett (Xylem/Sensus, NC) |
|---|---|---|
| Team size | 51-200 company (eng team ~50-100) | ~120 engineers |
| AI adoption | Claude Code 80-85%, production, 50%+ | Copilot corporate + Claude Code organic |
| Primary pain | Code review failure, cross-codebase context loss, rigid guardrails | Production bug from AI-generated code, no postmortem capability |
| Concrete incident | No single incident — systemic failure pattern with code review tools | Null-check bug: agent removed safety check AND its test |
| Purchase intent | 6 active initiatives, Jun 2026 target, 23 completed meetings | Pre-budget, 0 initiatives, 6-9 month window |
| Close strength | ”I’m intrigued” (soft) | “Let me talk to my senior engineers” (specific) |
| Archetype | 1 (Engineering) + latent 2 (co-creation with regulated clients) | 1 (Engineering) + latent 2 (utilities/metering) |
| Call quality | Strong discovery, pitch too compressed | Good discovery, talked too much |
Kevin is the stronger buyer. Brett has the better story (null-check bug), but Kevin has the purchasing behavior, the timeline pressure, and the more sophisticated understanding of the problem space. Kevin has evaluated and rejected multiple tools. Brett hasn’t looked yet.
Follow-Up Actions
- Send Sagetap follow-up message — see draft below
- Get Kevin’s LinkedIn — he didn’t share on the call. Send via Sagetap message or find manually
- Prepare demo tailored to Kevin’s pain points: code review failure, cross-codebase duplication, the “two doors” triage concept
- Add Kevin to the main Sagetap evaluation note with post-call update
- Create a person note for Kevin once real name and LinkedIn are obtained
- Connect Kevin’s “two doors” concept to the annotation pipeline roadmap — this validates a specific product direction
What This Call Adds to the Broader Discovery
New pattern: the code review trust collapse loop. No other discovery call has described this cycle so clearly. Tools get evaluated → tools fail → engineers lose trust → manual review resumes → manual can’t scale → repeat. Pathbase’s differentiation: it’s not another AI code review tool that will enter this same trust loop. It’s the data layer that makes review (human or AI) actually work by providing process context, not just output analysis.
New segment signal: professional services / co-creation firms. Kevin’s company co-creates with enterprise clients inside their codebases. This means code quality isn’t just an internal concern — it’s a client deliverable. Bad agent-produced code doesn’t just create internal bugs; it damages client relationships. The stakes are higher than for a pure product company. Other consultancies and co-creation firms (ThoughtWorks, Slalom, etc.) likely have the same dynamic.
Jellyfish as anti-pattern. Kevin’s dismissal of engineering intelligence tools as “rigid” and requiring heavy tagging/configuration is a competitive positioning gift. Pathbase learns from the trace — it doesn’t require engineers to annotate their work upfront. “Not another Jellyfish” is a compelling negative positioning for engineering leaders who’ve been burned by those tools.
Related
- 2026-05-14-sagetap-inbound-evaluation — Full Sagetap batch evaluation (Kevin = Priority #1)
- 2026-05-18-brett-1949-sagetap-call — Brett Dolecheck call for comparison
- 2026-05-10-sagetap-onboarding — Campaign setup
- Pathbase — Product context
- discovery-sprint-tracker — Prior discovery signal
Transcript
Bryan 0:00 Pull up my notes for myself in a second, but you’re based in the UK, right?
Kevin 0:04 Yeah, at the moment I’m in actually Dublin Ireland, but yeah, this part of the world, yeah, various airports and buildings and other things.
Bryan 0:16 Awesome, awesome. And how long have you been on Sage Shop,
Kevin 0:22 I few years now tend to be spiky on, so sometimes I’m very active, and other times it’s no busy, gotta do work, so there’s a lot of interest, and then this one was I’m probably in one of my quiet periods right now, but I saw this opportunity come up, and it’s an area, as I replied in the responses, that we’re very keen on, as we haven’t found anything that was really worked for us so far.
Bryan 0:57 Awesome, that’s great to hear, and I’m relatively new to the platform. This is my second stage call so far, but the answers that you submitted for the campaign that I posted were very exciting, and like the kind of the kind of entry point to the conversations that I’ve been looking to have. I guess this is a bit meta, but like I’m using AI myself to sort of review a lot of these responses, and like kind of formulating in our knowledge base the types of questions that we want to be asking, and Claude was very excited by your answers. You were the most exciting answers to your initiatives, and your answers were the most exciting to Claude.
Kevin 1:41 Actually, I didn’t does that mean we are the most like screwed up thing?
Bryan 1:47 No, not at all. I mean, like it’s not so much about like who has the most pain, but I think who has like the best grasp on like awareness of.. I think everyone has this pain, right? I’ve talked to you, know the even like coming to like Stripe, right, talking to like a senior engineer there, like they have these problems, are just very pernicious in this current paradigm. So I think having a grasp on it and trying to understand is important, and so yeah, that’s I think you’re coming at it from the right perspective, from my perspective. So you mentioned that using Cloud Code Cursor Copilot, and using it in production, and around 50% production work, and then R and D prototype projects and and that’s that’s also pretty much the same since, since last week.
Kevin 2:47 Yeah, public code would be the winner by far. Yeah, vast majority are using that.
Bryan 2:56 Yeah,
Kevin 2:56 Copilot tends to get used in co-creating stuff with clients. Yeah, it’s to be a big Microsoft worldview. Yeah, even though it’s just not as strong and can be frustrating for all sorts of other reasons. Got a few people using Codex, but hasn’t really taken off in the way that’s cloud code is probably 80 85% of the work is going through phone code,
Bryan 3:31 yeah, that makes a lot of sense. There was the huge inflection point, you know, end of last year when four or five hit, and like it just suddenly, it was like a set function change in terms of the ability to get to an outcome without a lot of a lot of direction and structure like it was a year ago. How are you? What is it like your code review process look like now with half of your production work going through or being generated even partly or in a hole by AI agents.
Kevin 4:03 I think it’s changed a lot in that it would have been before very traditional and the pace of reviews coming in were manageable for the senior guys to provide adequate review time, and sometimes they would do it in even in kind of peer review sessions, where they would actually review with the person, and so on. Now we’re seeing that the volume of, you know, PRs coming through is so much higher, especially when people turn on kind of full agent mode. It’s under the peers are bigger, you know, so there’s, there’s a lot more, you know, there’s, there’s potentially hundreds of commits in them, and you know, you’re, you’re seeing that that burden of, from my perspective, twofold, one either it becomes a giant bottleneck, or people skip them, because you’re just like, oh
Speaker 2 4:59 yeah, this, it’s probably okay,
Bryan 5:02 yeah,
Kevin 5:02 like, and we’re seeing that, and you know, and it’s essentially pushed the problem just further down, and there are, like, we have reviewed other tools, you know, we’ve been incredibly disappointed that we have industrial GitHub, but we’ve been incredibly disappointed with just its code review capabilities. It’s either one of the complaints I get, and I’ve seen this myself, you know, where you’re moving really fast, and you go in, you get it to, you know, do a review, like, you know, I’m, you’re holding me up now, and you know, and I want to, I want to keep moving fast, so it misses a lot of stuff. Yeah, so like we’ve, we’ve been one of the one of the things we do in our organization, we’ve got with a very healthy QA practice, and you know, we use QA for a lot of our well automated testing and review cycles, and so on, and we’ve been, we’ve been using those techniques to test the code review, you know, how good is it at code review reviewing by purposely putting in really bad stuff into the code to see if it’s easy to highlight, if it’s easy to pick it up, is it following, you know, is following basic stuff that we might, some of our clients would be very, like, you know, professional services areas. We would be doing a lot of co-creation work, and they’re very particular on the structure of the code.
Bryan 6:29 Yeah,
Kevin 6:29 naming conventions and stuff like that. So, you know, we, even though it’s not the best to be using it like a linter, you know, it’s it’s sometimes used like that, you know, as just a really quick way of just finding out, you know, what way you know the code base is structured, and they are, you know, files are way too big, yeah, the basic stuff, not the complicated, but what we do find is where most of them struggle, and this is this is probably the big area, if for instance we’re working on, you know, engineer A is working on something over here, and that’s fine, but they’re in full agentic mode, and a lot of the time we’re still in a place where the bots, cloud code, or any others, even if they read all the code, they still miss, you know, they’ll rewrite a function that’s already there over here somewhere, or they’ll duplicate something, or they’ll miss if we’ve got some part of the UI where we’re upgrading a workflow or some new business rule put in, but that similar functionality is elsewhere in the code base, and it’s not picked up, you know, and, and the engineer didn’t pick it up, you know, the agent didn’t pick it up, you know, and then it gets missed during the code review cycle, and we might find it in, in a QA, you know, automated testing from a QA perspective, it’s like, how do we miss this all the way through? Yeah, and you know, and, and from what has ended up happening is those tools either end up losing credibility in the engineering community, because the guys will go, okay, it’s just, it’s not helping me, you know. It means now what I’m doing is I’m getting the tool to review something, now I have to renew the output of the tool. It’s actually, why am I doing that? Why don’t you spend the time to review the PR myself, you know, so that’s that’s what we’ve ended up seeing with a lot of the, a lot of the different tools that we’ve been reviewing over the last while, they just, they have the headline, you know, kind of marketing, you know, an AI code review is going to help you do this is great, and we really need it, and we’re not seeing demos seeing that
Bryan 8:48 it demos so well, right? At first, right, it’s like they’re very good at, they mean like LMS are very good at producing seemingly good output, and if you’re looking at a demo, or you don’t intimately understand the code base, and like the edge cases, it’s hard to realize the shortcomings.
Kevin 9:05 Yeah, and we are seeing it, like, if you’ve got a traditional three tier, you know, environment, you know, and by that I mean you’ve got some storage there, and you know, broadly speaking, you know that they’re it, they’re better set up for that, but if you’re in agentic mode, there’s a lot of nuances around the way the code is set up, that it’s, it’s particularly what I find is if you’re in, so if I’m building an application that’s interactive, that a human is coming in, you know. There can be error messages, there can be other interact, you know, things that are going to come up if I’ve missed something, perhaps in the code. So, if the error, if the developer missed something in the feature, or something has happened, it would probably throw an error, an exception will probably pick it up somewhere. But if I’m in agentic mode, the whole idea is that these are fully automated, you know, and that you’re running something through, but you’re, you’re also trying to build in the logic, essentially the non-deterministic part of a lot of the logic, where, depending on what’s actually happened in previous steps, you know, if the, if the agent is is trying to interpret something and it comes back with an unanticipated direction.
Bryan 10:24 Yeah,
Kevin 10:25 are you picking those things up?
Bryan 10:27 Yeah,
Kevin 10:27 you know, are we? Are we picking things up, and particularly if are we picking something up where the agent is getting elevated privileges, or are we, you know, doing something that it wasn’t asked to do, you know, that you know, and for us, the way we look to do that, we’re trying to pick up, in particular, the guardrails all the time, like, so for us, a really, really important part of our QA, when it’s in agentic mode, is are the guardrails, you know, like, you know, security or other things, you know, to, you know, prompt injection and things like that, so if it’s getting an input and the prompt injections can come from anywhere, so if it’s even if it’s something simple, go go read a bunch of files in a code somewhere and do something as part of a process, there could be some injection in the files that it’s reading that steers it off in a direction, and we want to be able to pick that stuff up, you know, to prevent it. I know we often pick up everything, but the idea with the code review tools worked when it checked that the logic in the code is actually picking that stuff up, so that we don’t, you know, have to continue. Like, like the idea is that the code review tool should be strong enough to recognize that there’s a weakness in the architecture of what you, you’ve built in to highlight this, because you should be better at picking up a really complicated code base with lots of things going on, you know, and highlight to whoever is reviewing it that, hey, look, you need to look over here, this doesn’t look right,
Bryan 11:59 yeah,
Kevin 12:00 and and we’ve seen that’s been missed a lot, so what ends up happening is we ended up just a normal way.
Bryan 12:07 Yeah,
Kevin 12:08 so
Bryan 12:08 yeah, that’s kind of like, yeah, I think that’s something that my co-founder, CTO, he worked at Google for a long time, you know, kind of coming into a startup world with a lot of perspective on how large organizations should do stuff at scale, right. Google is typically an example of, like, okay, infinite resources and, like, a lot of smart people. How would you do it, right? So, it may not be achievable for a typical organization, but it’s a good example of, like, what you should be trending towards. And, yeah, he’s been interesting, like, he was a little hesitant to get into AI, like to like agentic coding, I was a bit more, as somebody who’s not as good a program as he is, like I was kind of excited a year ago to jump into cloud code, but so it’s been interesting to see him on this journey and seeing a lot of things that you were just mentioning as shortcomings or pain points around it, and so he’s been creating kind of like different open source tooling to kind of support that along the way internally of like, how do you kind of have these anchor points in the code base, and how do you sort of bake understanding and the contextual relevance into, you know, like not having things in different silos, right? You have your Git, and then you have your agent sessions, and you have your durability, and when they’re siloed, you don’t, you don’t get the benefits there,
Kevin 13:21 and we’re also trying to react to things that we’re not creating such a rigid structure. This was something that really struck me, and I can’t remember sending out a tool that we use to keep it it’s in the overview, a YC company, I remember the exact,
Bryan 13:50 okay, it’s okay, a
Kevin 13:52 YC, but I think he’s a Cuban, something like that, but I remember we were looking at it, and we started to build in rules and parameters around, you know, like that, so the agents themselves using, say, a version of Claudine’s doing the same, it’s fine, and so we’re we’re building an extra parameters rules around that to compensate for where it’s the boundaries are, what the model is and isn’t doing, and what it’s good at and what it isn’t good at,
Bryan 14:19 yeah,
Kevin 14:19 but then the model starts getting updated, yeah. Now all our rules are wrong, yeah. So we’re now building rules in that are too restrictive for the new models, yeah. You know, and so on, and sometimes that doesn’t matter because you’re picking them up, but sometimes you’re going to be overly, you know, critical. But we just want to have, you know, when we’re building in tooling, we want something that is flexible, and we use, you know, we’re using code reviews for even reviewing things that are outside of traditional code reviews, and you know, and what I mean by that is, we have a team, we call it the Digital Performance Team, and they keep an eye on the business perspective about how real users use real products in the real world, and we’re sort of picking up everything from, you know, Google Analytics feeds to other signals in the products, you know, with products that are, you know, assessing everything, so the checkout funnel are, you know, everything, and they’re not only just checking for things like doing A/B tests and all that sort of normal stuff, but they’re also picking up things where users are getting frustrated in the process, and so on. So we’re starting to, you know, but they’re building in code in those processes to check those things that are kind of running outside of the standard, you know, the system itself, you know, it’s something we’re layering in, and we want to make sure that that stuff isn’t going rogue as well, you know, so those guys are using code in a different way, yeah, and using third party tools and using kicking into MCP and all sorts of stuff, and I would like that to be covered as much as possible with review tools, because traditionally it wasn’t. They weren’t applying any development techniques or processes or quality or anything. It was just been done by let’s just hack to this thing to death and hope it works. We want to try and bring some rigor in, so it will be great to apply all of the tooling that we typically do from a developer, but this is to a community that aren’t heavy engineers, they’re coming out from a slightly, they’re technical, not not not deep coders, but but now we’re in this world where using, they’re using all these tools, our designers as well, like our designers are using, you know, they’re pushing into writing the very first versions of a lot of new features, you know, with code, and we want to start getting familiar with GitHub, or getting familiar with checking code, and just, you know, we’re trying to introduce more rigor into the process, yeah, because if we can push that upstream rigor into code, it means when we start going into full production mode, with our engineering squads to take it on, where we let’s rework,
Bryan 17:07 yeah, and
Kevin 17:08 that’s where the tools can have decently a new generation of tools, but I haven’t found some yet,
Bryan 17:14 yeah,
Kevin 17:14 you guys will,
Bryan 17:15 yeah, there’s a lot in there that makes a lot of sense, I think you seem like a pretty, like you’re pretty far into this journey with a team at scale working with with agentic coding, and yeah, the there’s like a certain amount of lived experience with having built code before using AI to write code that is very critical, and like the separate, it’s like subtly critical, how important that that is, because if you don’t know what good is, and if you don’t understand the reasons that things are done a certain way, you can get really far and think that everything’s great, and then slowly you’re kind of just in this like descending into a black hole of LM psychosis and isolation on whatever you’re building, not realizing how this fits into a broad organization, sort of Chesterton Spence, of like understanding why a thing was there before you rip it out and throw it away, and I think that’s kind of a phase that a lot of organizations, or a lot of tech forward organizations, are now where engineers have like seen the sort of like the initial kind of utopia of this and then realizing like you know the dark underbelly of all this and again kind of put structures in place to try to get the most of the upside and minimize the downside and this new phase is about how do you share that like bleeding edge learning with the rest of the organization who are and should be working with agentic coding, but don’t have that lived experience, both prior to AI and in the initial rollout of AI. So, I think you’re right in the middle of, like, the hardest problem right now.
Kevin 18:53 Yeah, like that will be one. It probably wouldn’t be at the top of my list, but it’s just like, if we’re going to go and do it, we might as well to other stakeholders, other users around the organization, but I think for me one of the big areas that we’re focused on is, and we’ve been from an automation, I’ll describe it as an automation, you know, we have use cases where I’ll describe it in your typical kind of maintenance or business as usual teams, where you know you’re making a lot of small changes, the continuous tweaking process is constantly on, and instead of every single thing having to be PRed by just the team and every single line of code written, we have a pyramid of changes that some are going through fully automated, where we’re not stopping the PR for human review, we’re letting it go through, yeah, you know, but you know, and then there’s others, we’re, you know, full, yeah, humans in the loop, but the trick that we’re trying to get to, we’re sorry, I should rephrase that, we would like to let some of those PRs through because they’re non-contentious changes that are going through that the agent has executed perfectly.
Bryan 20:09 Yeah,
Kevin 20:09 how do we assess that?
Bryan 20:11 Yeah, it’s sort of like a radio and previous kind of thing, for yeah, yeah. And the
Kevin 20:16 reason we want to do it is to remove noise out of the system that we have, you know, it’s it’s taking a lot of engineering time for many of these, you know, like just a volume of changes that are coming through, or being requested, it could be from the aforementioned team that I mentioned, doing all the digital performance stuff, and it’s just kind of small tweaks, and you know, let’s, let’s try this thing with a green button and see what it’s like, so it’s not a massive change, it’s a, it’s a one parameter in that, you know, that’s going through, so we know it’s going to be executed pretty easy,
Bryan 20:51 yeah,
Kevin 20:53 but we now have this whole infrastructure of everybody having to review everything, spelling it on land, so there’s a certain category, so a future level of tools that can, that can be configured in a way to say there’s two doors, you know, if you meet a certain bar, just fly to PR, just go full, full automated, you know, otherwise go for human review and door B, that would be brilliant.
Bryan 21:19 Yeah, I think understanding, like, what are the sort of like linchpins in this particular code base, and like, if you’re touching these, like, this is automatically like needs double manual review, whereas, like, okay, you’re changing documentation, and some of these things are intuitive, but as code bases are getting larger and everyone’s moving faster, there’s a lot of, I guess, intrinsic, like, I just like ambient shared awareness of things that were kind of changing, or like PRs. You might see your coworkers’ PR and be like, ‘Oh, like they’re touching that thing that I worked on, let me tell them that thing. But everyone was just like full on with, like, you know, eight to 10 clogged code sessions live, and a lot of the things that we took for granted are no longer present in the way that we produce code, which is again another sort of symptom of the times that we’re all adjusting to. I have a lot of ideas on, and obviously like thoughts on how to do this better. And yes, we have like a platform that we’re launching and looking for, kind of like early early customers to onboard with this. So, yeah, at a high level, everything you’re saying makes a ton of sense, and you’re preaching to the choir with, with, with, with these problems, and yeah, I, you know, as some of it uses LMS a lot to like Saturday check things and parse lots of data, like it’s very alluring that, like, okay, like you feel like you’re getting something really quickly, and I think, especially with code review, you can feel like your product, then realize, like, oh, I just wasted like all this time, and I’ve got to go back and manual for everything anyway, because what it was giving me was just sort of like trying to make me feel good, not it wasn’t actually being critical or like reasoning about anything, it was just, you know, being sycophantic, so
Bryan 22:57 yeah,
Bryan 22:57 trust
Kevin 22:58 is the word that I use, yeah, it’s suddenly you end up in real trusted,
Bryan 23:02 yeah,
Kevin 23:03 and even when it starts getting right yourself, is it correct?
Bryan 23:08 Yeah, yeah,
Kevin 23:09 you end up manually reviewing anyway.
Bryan 23:11 Yeah, so everything that we’re building right now is optimized for manual human review. We’re kind of taking, like, taking it back to first principles approach and looking for essentially like novel data structures that create better, you know, better information presenting for the human to make better decisions. I think in that process there’s a substrate on which to build a significantly better AI-assisted automated code review down the road, but I think it’s all those companies are just way out over their skis with trying to just throw it out, throw it at an LLM and get an answer, like it’s it’s like a friend of mine called it the accordion problem of like everyone’s using like LM to generate things, and there’s like it’s too much for a human to actually consume and understand, so that a human has to take it and, like, like use an LM to press the accordion back together, just to have an input, and you have this whole.. it’s like playing like a game of telephone with information and code, and it’s a very weird, very weird time. So, yeah, I think the I alluded to before, this idea of kind of creating anchor points, or you know, qualification points within the code base, that’s an open source thing I can send you that my co-founder is working on, and that’s that’s sort of like put a pin in it as like a product feature. The really core thing that we’re working on right now is also an open format that we’re calling Tool Path, which is a like a structured provenance format for both code and the the process that went into it, which again, if you’re in an agentic coding paradigm, is the prompts, but it’s also the tool calls and the harnesses that are writing the code are getting increasingly complex, leaving, you know, state and artifacts in different places on your laptop, or whatever, maybe, and so we’re basically collecting all that and creating like a super graph of the both like the code versioning and the actual process that went into it, and then as we get further in this, the ability to kind of like further annotate that graph, and by having this all in one place in a structure that a human can kind of review it, there’s a whole number of interesting things that we do. We have four minutes left in the call, so I can’t regale you with all the cool things that we can do with that. But essentially, like when you open a PR, you can have a link and you can click that link and you can see that every prompt like interlinked with the entire process, and a lot of like get code review is very focused on like the end output, right? Okay, here’s, here’s, here’s the steps of the code that went into it, but you don’t see the dead end, right? You don’t see the half a day that an engineer saw, like, going down a certain path, and they’re like, ‘Oh no, this is a dead end, I have to throw this away and go back, whatever I was doing. And I think that is the kind of thing that might have been caught in, like, you know, stand-up or conversations, but in this world, like, everything, everyone is just going full throttle, all that really important context is just being kind of thrown away, and so our primary goal is to capture as much context as possible and structure it in a way that it is consumable both by humans and by the coding agents themselves, right? That’s that’s the high level, and so that’s an open format, and then we have like a hosted platform, sort of like a GitHub for this this hosted format, so you can again have like those links to click into from a PR and see that full graph. You can on the roadmap is being able to, you can sort of like boot up an agent, kind of like almost like taking out a cadaver and being like, okay, you’re about to commit a grievous error here. Can you explain to me, like, why you would do this thing, sort of like little Minority Report style? But yeah, I think it’s, it’s very, very needed from everyone that we’re speaking with. This, this type of data structure,
Kevin 26:57 that’s actually interesting, you know, because I’ve always been disappointed over the years with tools in kind of an adjacent space, I’m going to call it engineering intelligence, and stuff like companies like Jellyfish and others like that that are promising all of this stuff, and you end up then having to tag and procure your code base in such a way that becomes again a very rigid thing, and you know, and for me, the way you’re talking about is, if you know, just that example that he gave, that there’s, there’s, there’s organizational and engineering knowledge that kind of went down this path and got bulk and come back again, yeah, you know, you don’t want the next developer to go do the same thing, and you know, and continue to repeat, and everybody discovers the same problem all the time, yeah, you know, and it’s, it’s, there’s, there’s great knowledge in that, in just learning the code base and getting people up to speed, and not only is it important to me, I think it’s, it’s, and I’ll leave you at this point, you know, which I think, if you guys have something in this area, it’d be very interesting that the agents are going to be good at reading all the code, but I also want my engineers to be really good at looking at the breadth of the code base, so that they can spot the things where they were the agents have gone wrong, or they’ve missed something, like the oh, and there’s a function that does that over there, or there’s a class, or there’s some module that already does that, or we have a whole UI built on the admin system that does exactly the same thing that you’re planning to rebuild on the user system or something. Yeah, you know, why don’t we just combine the two, and the agent might miss that, whereas the human is aware, and anything that we can, you know, if you apply the same to Business Logic on 50 other things, the same thing applies. We’ve got these engineers that now have this breadth of the whole entire code base that can really guide the agents in this co-creative mode, that would be fantastic.
Bryan 28:45 Yeah, that’s amazing. And I love the way that you described it back to me, that’s fantastic. Yeah, so we’re at time. I would love to continue the conversation. I can loop in my co-founder to kind of bring in some more of the deep technical perspective, and we can kind of dive right into a demo and kind of walking you through some of the stuff that we’re building, if you’d be up for that.
Kevin 29:05 Yeah, and when you, when you’re launching, using months or weeks,
Bryan 29:08 like any day now, right? There’s just some little, little ironing out that we want to make sure, because we don’t want people to have a bad experience the first time they use the hosted platform. So, yeah, it should be, it should be any day now. It’s usable now, right? We can take you through, and everything, and if you’re like, I need this today, we can probably get it live for you today. But yeah,
Kevin 29:30 look, it sounds good, Bryan, and
Bryan 29:32 awesome.
Kevin 29:33 Yeah, we’ll see about seeing taking it forward, and yeah, it’s always.. it’s always one of those weird times of the year, as well.
Bryan 29:43 Yeah,
Kevin 29:43 I’m intrigued. Put it this way,
Bryan 29:46 would you like me to add you on LinkedIn while we’re here?
Kevin 29:51 Yeah, let me.. I’ll send you the details after.
Bryan 29:54 Okay,
Bryan 29:54 perfect. Yes, that sounds good. Well, lovely chatting with you. Have an awesome day, and looking forward to talking again soon. Have a good one. Cheers. Bye.