Everyone Can Run Agents Now: the Edge Is the Opinion You Hand Them

Everyone can orchestrate agents now, which makes orchestration the floor; what stays scarce is the conviction you hand the agent.

8 min readBy Matthew Stublefield
Silhouette of three women running on grey concrete road

I got sent to a software company for two days of Jira and QA training, and by lunch on day one it was clear that training wasn't the problem. What I did next is the mistake the product world is about to make with agents, only much faster and much wider.

One shared QA team served several product groups, and that team had become the flashpoint for every conflict in the building. Competing priorities collided there. Blame landed there. My instinct was the one I've had my whole career, the same instinct that has me rewiring a light switch at ten at night instead of calling an electrician: I can build this, I can fix this. So I built them the coordination system, intake and queues and ownership and the rest of it, and then I flew home pleased with myself.

It didn't hold. The systems weren't wrong, and that org genuinely did need a better way to coordinate a shared team. What they lacked was any practice at thinking through their own systems, their own communication, their own conflicts, and I hadn't given them any of that practice. Two days buys a build. It doesn't buy the judgment to extend the thing, argue with it, or rebuild it when the org shifts underneath it, and that's closer to three months of work alongside people than two days of work on their behalf.

I've come back to that engagement a lot this year, because the trade I made is now being made everywhere at once. The 2026 consensus is that a product manager's job is now orchestrating a fleet of agents. Every product blog, bootcamp and tooling vendor has published some version of it, and they converge on identical advice: run the fleet, invest in judgment, keep a human in the loop. The consensus is right, and it describes the new floor rather than the new edge. When everyone can spin up the fleet, running it is table stakes, and what stays scarce is the conviction you hand the agent. The durable form of that conviction is a knowledge base with a judgment layer on it, which is exactly the thing I didn't leave behind on that engagement.

AI took the extraction work, not the judgment

As one 2026 survey of the shift put it, four product-management workflows moved from human hours to agent minutes: drafting PRDs and documentation, synthesizing customer research, running data analysis and reporting, and generating feature ideas. Twenty solution directions against a well-framed problem is one prompt now, so the scarce skill moved to framing the problem and throwing out nineteen of the twenty.

My own calendar shows this in a way I didn't expect. Getting from an initial idea to a fully developed epic with stories takes me about as long as it ever did. The research inside that time collapsed, so a much larger share of the same hours goes to thinking, and my guesses are better educated than they were. The hours didn't shrink; what they're spent on did. I've written about what happened to the PRD as this played out, and about intent as the constraint on the engineering side of the same shift.

Agent fluency has already become a hiring filter

Orchestration got priced as an entry condition before most teams learned to do it. Dexity's analysis of 654 live PM job descriptions found that 28% explicitly call for agents and 85% mention AI or ML. That's Dexity's own dataset, published on a page selling their training program and not audited by anyone outside it, so read it as a directional signal rather than a finding. It matches what I see in the postings that cross my desk, which is the most I'd claim for it.

The other half of the picture is that almost nobody is good at this yet. McKinsey research puts 23% of organizations scaling agents in at least one function and 39% still experimenting, with 10% or fewer at scale in any single function. I'm two hops from the primary on those numbers, since they reached me through a vendor page citing Forbes citing McKinsey, so hold them loosely. Directionally they say most agent deployment is still pilot-stage, which is a strange thing to sit next to a hiring filter.

Both things are true at once. A filter is something you clear, and clearing it gets you into the room. It doesn't win the room. If you're competing on the thing that 28% of job descriptions now require, you're competing on the requirement rather than on anything of yours. Models synthesize, distill and repeat extremely well. They cannot supply the point of view they're handed, and that's the part you own. It's the same move I described when knowledge stopped being the moat, one layer further up.

A shared knowledge base is a search engine until you add the judgment layer

Pipe your call transcripts, support tickets and sales notes into a vector store and you've built a search engine. A good one, probably. It'll answer "what did enterprise customers say about SSO last quarter" in seconds, which is genuinely useful and not an asset, because any competitor with the same tickets and the same tooling gets the same answers.

The judgment layer is what makes it yours. Three parts, in order: what they said, what they meant, and then what you think given where the product needs to go. The first part is a recording. The second is interpretation, which is where most product judgment actually lives, because customers describe symptoms and ask for features and only sometimes name the problem. The third part is the one nobody writes down, and it's the only part a model can't generate for you.

Build that and something specific changes. An engineer working a ticket at 2am, who would otherwise ping you and wait until morning, can ask the system instead and get a glimmer of the call you'd have made. Not your exact decision. A defensible approximation of it, with the reasoning attached, which is enough to keep moving and enough to know when to stop and wait for you. The retrieval quality barely matters here, which is why I keep arguing that the retrieval problem is usually a corpus problem wearing a technical costume.

I've built a handful of small software tools in the past year. I'm not an engineer, and no agent made me one. What changed is that twenty years of written-down opinions finally have somewhere to go, and the machine is a good deal more useful when it's arguing from my recorded positions than when it's arguing from its own inference.

Deciding what to hand an agent

Delegation gets easier once you stop asking whether the agent is smart enough. A delegation model taught by Meta's Sanaz Alexander, a Director of Product Management there, maps a task on three axes: complexity, frequency and reversibility. High-frequency, low-complexity, high-reversibility work suits an agent. Low-frequency, high-complexity, low-reversibility work stays with a human. It's a clean framework, and the reason it holds up is that good managers already use it on people.

Reversibility is the axis people skip, and it's the one that matters most. Competitive summaries, release notes, first-pass research synthesis, the tenth variation of an onboarding flow: all cheap to be wrong about, all easily undone. Pricing changes, deprecations, a public commitment to a segment, the architectural bet that shapes the next two years: wrong there is expensive and slow to unwind. An agent's answer arrives fast, documented, and confident, with nobody accountable for it. That combination is the Cheap Yes, and its endpoint is the Expensive No that someone else pays for later.

What to build this quarter that compounds

Build the judgment layer, not the prompt library. Prompt libraries are the artifact everyone reaches for, because they're easy to produce and they feel like infrastructure. They also go stale with every model release. The record of what your team decided and why gets more valuable the longer it runs.

The practice is smaller than it sounds. After each customer conversation, three lines go into the shared base alongside the transcript: what they said, what they meant, what we think given where the product is going. Same for a lost deal, a support escalation, a spike a designer runs. It's a couple of minutes of writing per event, done by whoever was in the room, and it's the whole difference between a corpus that answers questions and a corpus that carries a position.

The alternative is what most agent programs are measuring right now, which is thinner than it looks. Ten prototypes in a week tells you the agents ran. It doesn't tell you that a single decision changed, and activity counts that don't resolve into decisions are the same failure mode I described as prioritization theater, running at a much higher clock speed. If your agent program can't point at a decision that went differently because of it, you have throughput and not much else.

This is most of what the Intelligence stage of my own work turns out to be in practice. Research answers one hard question once. Intelligence keeps answering it as the world moves, which requires exactly this kind of maintained corpus with a recorded point of view on it. Strategy turns the accumulated evidence into a direction somebody is willing to put their name on.

The conversation that started me down this road is worth an hour on its own: Kevin Yien of Stripe, interviewed on The Skip, talking with Nikhyl Singhal about what comes after "become a builder" for product managers. Yien runs merchant experiences at Stripe, which covers the dashboard, the mobile apps and the agentic Console they announced at Sessions, so he's shipping into exactly this question rather than theorizing about it.

I could have taught that company to run their own coordination system, and I'd guess three months would have done it. I built it in two days instead, because building was faster and it felt like help, and it left them with an artifact nobody in the building could argue with or repair. Agents make the building faster still, so the temptation is much larger now and the miss is identical.

If you're working out what that judgment layer should look like for your product, email matthew@fieldway.org to discuss.

Sources

Want help running a sharper practice?

The reading and synthesis behind your client work, handled – a living deliverable kept current, so more of your time goes where your name is actually on the line.

See how this works for advisors