Executive Overview
For roughly a decade, the role of the Chief Product Officer (CPO) in B2B software was remarkably straightforward. You brought your coffee to the morning stand-up, outlined the features slated for the annual release cycle, quietly bumped a few complex items to the next quarter, and—when an impatient investor asked about a missing capability—reassured them that it was "on the roadmap." Sometime later, it usually arrived.
That predictable era officially ended with the close of 2024.
Today, every B2B product leader is under intense pressure to ship functional AI agents ASAP—agents that customers are actually willing to pay for. The stakes are immense: when companies successfully monetize their AI offerings, the market rewards them immediately. (Atlassian, for instance, finally cracked its AI monetization model, sending its stock soaring roughly a third in a single trading session). Most of the industry, however, remains in the grueling trenches of figuring out how to do the same.
To cut through the noise of traditional industry events, SaaStr banned standard presentation panels for its SaaStr AI event, viewing them as inherently unengaging. The sole exception is when the leaders on stage share a deep, operational rapport. To explore the reality of shipping AI agents into high-stakes environments, SaaStr gathered four elite CPOs: Anneka Gupta of Rubrik, Emrecan Dogan of Glean, Anique Drumright of Harvey, and Rachel Wolan (then CPO at Webflow).
Selling to cybersecurity teams that cannot tolerate a single erroneous action (Rubrik), powering the underlying context layer for enterprise agents (Glean), or servicing high-powered law firm partners who fund software out of their own pockets (Harvey), these leaders face vastly different buyers united by a single, formidable problem: how to build, deploy, and trust autonomous software agents.
Detailed Chronology: The Evolution of the B2B Product Roadmap
The transformation of the CPO role from feature-scheduler to agent-orchestrator has unfolded in rapid, high-pressure phases over the last 24 months.
Phase 1: The Illusion of the Simple RAG Feature (2023–Early 2024)
Initially, enterprise AI was approached as a retrieval-augmented generation (RAG) problem. Companies ingested setup manuals and troubleshooting documentation, allowing users to query a chat interface. The resulting answers were, as Rubrik’s Anneka Gupta described them, merely "okay answers based on whatever was in the documentation." Product teams treated AI as a wrapper—a nice-to-have chatbot bolted onto the side of legacy dashboards.
Phase 2: The Agentic Pivot and Scope Explosion (Late 2024)
As customers demanded actual workflow execution rather than just text retrieval, product roadmaps fractured. Building an agentic workflow turned out not to be a feature update, but a second, concurrent build of the entire software product. Companies quickly realized that scoping an agent roadmap like a minor feature release while pricing it like a full platform rewrite was a recipe for catastrophic delays.
Phase 3: Deterministic Execution and Guardrails (Present Day)
The industry has now entered a phase where probabilistic generation must meet deterministic execution. Leaders are learning that in mission-critical environments—such as cyber recovery or legal compliance—allowing a model to improvise its steps is unacceptable. The modern B2B product roadmap is defined by strict oversight: models propose and explain plans, but the execution layer remains rigidly controlled, auditable, and human-verified.
Supporting Context & Metrics: Insights from the Frontlines
The discussions at SaaStr AI yielded critical metrics and operational realities that every software executive must digest.
1. The Context Bottleneck
According to Glean’s Emrecan Dogan, if an enterprise user spends five hours a day interacting with an AI assistant, roughly half of that time is consumed by "building context"—feeding documents, updating memories, defining company-specific rules, and training writing styles. Retrieval is no longer the primary hurdle; performance is capped entirely by context. If your product requires users to paste the same background documents into every session, that setup time is the hidden tax on your product’s value—one that never registers in standard funnel metrics.
2. The Rise of Ecosystem-Native Usage
Glean’s growth trajectory highlights a fascinating shift in how enterprise tools are consumed. Glean now operates both as a standalone assistant and as a Model Context Protocol (MCP) server that external developer environments—such as Claude Code, Cursor, and Codex—tap into directly. Dogan noted that the aggregate usage of Glean embedded within these developer environments is growing faster than Glean’s native UI, proving that modern agents must meet users where they already code and work.
3. The Economics of Elite Verticals
When selling to elite law firms, Harvey navigates a unique economic barrier: software costs often come directly out of the partners’ own profit distributions rather than a standard corporate IT budget. Despite this hurdle—which makes enterprise sales exceptionally rigorous—Harvey has penetrated over 60% of the Am Law 100 and boasts more than 700 customers across 58 countries.
Official Statements & Core Takeaways
The CPOs on stage shared nine foundational truths regarding the reality of building enterprise agent architecture:
- Start with agents that cannot break anything: Rubrik’s initial agentic workflow focuses on forward-looking capacity planning—a task that takes humans a full day by hand. It reads, analyzes, and recommends, but never touches production. High-risk workflows come later, and only after human verification loops are established.
- Decouple generation from execution: In cyber recovery, Rubrik refuses to let probabilistic models improvise recovery steps. The model generates the plan, but the execution of data recovery is strictly deterministic and auditable.
- Abolish the central AI team: Rubrik deliberately avoided standing up a single, centralized AI team to build features for the rest of the company. Instead, they democratized the mindset, requiring every individual PM and engineering team to think agent-first.
- The friction of high-compute offline processing: While runtime fetch protocols like MCP provide quick access to data, they are bounded by latency and search APIs. High-compute offline processing—connecting employees, historical projects, and acquired subsidiaries before a user ever asks a question—remains the true differentiator.
- Domain expertise beats generic internet skills: Deploying generic sales-analysis skills downloaded from the internet fails in enterprise settings. Glean’s sales agents leverage years of specific company context, frameworks, and institutional knowledge to propose accurate updates directly into Salesforce.
- Bespoke deployment models: Harvey scales its engineering prowess by deploying legal engineers—many of whom practiced law for 8 to 10 years—directly into customer workflows alongside software engineers, ensuring the AI encodes the exact standards of the firm.
- Evolving review workflows: Elite law firm partners are no longer just reviewing marked-up drafts; they are actively reviewing and editing the agent plans generated by associates before execution takes place.
- Ultimate accountability remains with the platform: When asked who bears responsibility if a customer’s agent breaks a system, Anneka Gupta was unambiguous: "At the end of the day we’re still responsible… It’s our responsibility to make it obvious what the right choices are."
- The expanding surface area of unintended actions: Because modern AI agents can execute actions across both headless environments and user interfaces, customers will inevitably use your product in ways you never planned or anticipated. Software vendors own the downstream outcomes regardless.
Future Outlook: The Hardest Job in B2B Tech
As the software industry charges deeper into the agentic era, the profile of the successful B2B product leader has fundamentally transformed.
The modern CPO can no longer rely on superficial feature roadmaps or vague promises of future AI integration. They must successfully balance four competing imperatives simultaneously:
- Deploying agents that take real, autonomous enterprise actions;
- Safeguarding products where a single wrong action can instantly terminate a high-value customer account;
- Engineering verification loops that earn absolute customer trust; and
- Accepting total downstream accountability when autonomous workflows deviate from expectations.
Security teams are pushing hard for the speed that only agents can provide, fully aware that attackers are operating at machine velocity. Yet, they simultaneously refuse to cede final operational control. Navigating this delicate tension between autonomy and accountability is precisely why managing a B2B product organization today has justifiably earned the title of the hardest job in technology.
