The Product Manager's AI Operating System: One Triage, Eleven Pipelines, and the Gates That Make AI Output Shippable

Most of the workspace is refusal rules. Around sixty skills, eleven pipelines, and more of the text given over to what the model may not do than to what it should produce. Here is what that buys, and what it does not.

Share
A machined graphite slab rests on concrete: one wide channel labelled TRIAGE divides into eleven milled channels ending at printed labels, P1 Meetings through P11 Measurement, while three…
The Product Manager's AI Operating System: One Triage, Eleven Pipelines, and the Gates That Make AI Output Shippable

Most of the workspace is refusal rules.

That's the first thing a careful reader notices in the operating instructions for the product-management and business-analysis workspace I run on Claude Code. Around sixty skills, eleven pipelines, and more of the text given over to what the model may not do than to what it should produce. It's also the part that took longest to get right, and the part that makes the output inspectable before it goes in front of a client.

If your PM function got AI tooling this year, you may already know the complaint from both sides, because it's the one this system was built to answer. Delivery says the requirements arrive faster but not sharper. Sales says the proposals look polished and nobody can tell which numbers are real. The instinct is to fix that with better prompts. Better prompts didn't fix it for me. What did was treating AI product management as a production line with a router at the front and gates at the back. The interesting design decisions turned out to be the negative ones: which pipelines don't get the interview, which facts the model isn't allowed to supply, which review nobody can skip.

The case that the PM role itself has to be redesigned, and what the governance ladder under it looks like, lives in the PM AI playbook; its business-analysis counterpart is the BA AI playbook. This piece is the machinery behind both arguments: what a request becomes, which of eleven pipelines owns it, what refuses to ship, and what a person still decides.

An AI operating system for product managers, as I've built it, is three things. A triage that routes every request to one primary pipeline out of eleven. Two execution modes: Auto approves the plan up front, Manual confirms each step. And a set of gates that run in both modes. Those gates are source traceability on every claim, an adversarial review of every client-facing final, and logical isolation per client. The routing and the gates are what make the output inspectable. They don't make model choice irrelevant, and they don't replace the person who decides what "done" means.

Every number in this article comes from the workspace's own artifacts: the operating instructions, the eleven pipeline files, the project logs, and the dashboard rows of a test run against a fictional client called SmartSaver. The kickoff transcript that seeded that run declares itself a synthetic fixture. What the run establishes is that the machinery works end to end, on a client whose outcomes nobody can claim. I'll say so again at the points where it would be tempting to imply otherwise.

The first job of the system is to refuse to guess what the job is

A request lands: "turn this brief into requirements for the pilot." Before the model reads a word of the brief, the instructions require the triage block at the top of the workspace's operating instructions to run, and its first step is classification. Every product management AI workflow in the system starts here, and the rule is strict. The request is classified into one primary pipeline, P1 through P11, into a cross-cutting recipe, or into a supporting job like standing up a client's design system. Only then does any research or writing begin.

Those eleven are the classic BA/PM artifacts, each keyed to trigger phrases: meetings (transcripts, agendas, stakeholder updates); market and competitor research; discovery through proposal; requirements, whether a BRS, a PRD or a one-page feature brief; UX discovery artifacts; prototypes; presentations; release notes; demo videos; roadmaps and prioritization; and measurement and experimentation. Around them sit cross-cutting recipes that aren't jobs on their own: diagrams, document formats, fact-checking, architecture decision records, audience-tailored briefings, the adversarial review, the grilling interview, and a brand-voice check on client-facing copy.

"One primary pipeline" is the precise phrasing, because composite asks are normal. A discovery engagement chains P1 into P2 into P3 into P4 and then into P6 or P7. The triage routes to the pipeline that owns the requested deliverable; that pipeline invokes the others as sub-steps, and each sub-step reads its own pipeline file as it starts. What the rule forbids is the model deciding on its own that a requirements request is "really" a roadmap request.

What remains of the triage is short, and it's all refusals in disguise. If the request is ambiguous, ask one question with the candidate pipelines as options, never guess. Identify the client, because every job belongs to exactly one, and never infer the client from whichever design system happens to be installed. Run the brand check. Decide new job or continuation of an existing one. Confirm the scope in a single line ("P4: BRS for SmartSaver from the discovery brief into the dated job folder"). Then read the pipeline's file.

That last step has its own enforcement line in the instructions: the file, not recollection, is the spec. Executing a pipeline whose file hasn't been read in the current session counts as a triage violation, and a continuation session re-reads the file before resuming. The rule exists because a model that has run P4 five times will happily run it a sixth time from memory, and that memory may not contain the step that was added yesterday. (The intake is bundled, too: two or three question rounds, never five. A triage that interrogates the user is just a slower way of guessing.)

The workspace's CLAUDE.md triage block: every incoming request is classified into exactly one of eleven pipelines, meetings through measurement, before any research or writing starts.

In operating-model terms this is the workflows-and-handoffs layer, and it's deliberately boring. Its job is to make sure the expensive machinery downstream is pointed at the right artifact.

Auto mode approves the plan, not the output

The execution mode is step zero of every pipeline, and the two modes split on a single question: who owns the taste calls. This is where most discussion of AI in product management goes wrong, because "hands-off" gets read as "ungated", and in this workspace those are different settings.

Auto runs the pipeline end to end on its stated defaults and presents the finished deliverable. Explicit parameters in the request override the defaults; defaults fill only what was left unstated. Choosing Auto is the plan approval, so the model asks no mid-run questions. What Auto does not do is skip the gates (the instructions call them objective; some are mechanical checks, the finals review is judgment). Source traceability, validation passes and brand fidelity are required to run exactly as they would in Manual, the test-run logs record them running, and a gate failure means stop and report. Not ask. Not ship past it. Stop.

Manual is the interactive lane: a full brief through structured questions, a proposed approach, an explicit wait for confirmation, a draft, iterations, then the final. Structure, tone, visual direction and depth are decided by the person in Manual; in Auto they're decided by the pipeline's defaults and labelled as such.

Two details from the SmartSaver run show what "labelled as such" means in practice. The P4 requirements job ran in Auto and hit three edge cases the inputs didn't resolve (answers insufficient to identify a product, a search that returns no qualifying offers, an invalid or implausible reference price). The log records that they were written into the document's Open Questions rather than resolved by an invented threshold. The P10 roadmap job, also Auto, scored seventeen items with RICE and marked every Reach and Impact value as an estimate with its basis stated, because there were no live users to measure. Auto decided the defaults; it didn't get to decide the facts.

A caveat belongs here, and it's the one a skeptical reviewer will raise first. "The gates never skip" describes what the instructions require and what the test-run logs show. It isn't a mechanical guarantee. A person retains exception authority (a P0 finding can be explicitly accepted), and a procedural control is only as strong as the habit of not routing around it. I treat the rule as a decision-rights statement: Auto transfers plan approval to the defaults and keeps the release decision with the gates and, ultimately, with the human reading the gate log.

Brand is applied per job and never stored globally

Client isolation is the rule that looks trivial in a folder diagram and turns out to be the one the tooling fights hardest. Every client can have a design-system skill of its own, generated from that client's brand materials, carrying a design spec, design tokens, real logo assets and usage rules. The brand check runs right after the client is identified, and it only matters when the job produces something visual: prototypes, decks, demo-video wrapper elements, diagrams, a styled proposal. Text-only jobs (meeting summaries, research reports, release notes, measurement artifacts) log "Brand check: N/A" and move on. Requirements and proposals are conditional: the moment either renders a diagram, the job is visual and the full check applies, and that decision is taken at triage rather than after the diagram exists.

Identity is confirmed by provenance. Before a design system is applied, its manifest is checked against the job's client folder, and a mismatch stops the job even in Auto. The model is never allowed to infer the client from whichever design system happens to be installed.

Diagram engines are the awkward case. One of the two engines stores brand globally by design. Its onboarding rewrites the shared style guide, its marker file only resolves at the project root, and its profile library lives outside the workspace where any project on the machine can read it. Each of the three persistence paths the skill documents would let one client's palette leak into the next client's diagrams, so all three are disabled. Brand is applied per job instead: the semantic colour roles (paper, ink, muted, accent, link) are read from the client's token file and substituted into the generated diagram only. Nothing shared is mutated; the log line reads "Diagram brand: per-job token substitution". A value the design system marks unknown is left at the engine's shipped default and noted beside the deliverable, never invented.

When no design system exists, the rule is one question with three honest answers: create it now (its own dated job), proceed unbranded (legitimate, logged, re-offered next time), or provide materials later. Never silently unbranded. The SmartSaver prototype ran unbranded on a one-off token sheet, which the pipeline explicitly permits for a test client.

One more caveat, because the word "isolation" invites it. A client folder is logical isolation enforced by instructions. Nothing technical prevents a tool from reading a sibling folder; the rule does, and the rule is written as absolute and logged on every job precisely because it has no mechanical backstop.

Who answers the status question?

Not the model's memory. That one design choice carries more weight than it looks, and the information layer it sits in is worth seeing whole.

All production work lives under one root, one folder per client, one dated folder per job, and each job folder has exactly three children. Input Data holds the sources the user provided and is read-only by convention. Processing Data holds everything generated while working, all of it disposable except one file. That file is project.md, the job's memory: the first session creates it, and every later session appends to it with a dated heading, its decisions, its gate outcomes and its brand-check line. Final Deliverables holds only what was approved. A revision never overwrites a prior final; it gets a version suffix, and the first version keeps its name. The folder date is the job's origin and is set once; a continuation keeps the original folder rather than starting a second one.

On top of that sits a dashboard per client: Planned, In Progress, Completed. One row schema covers all three: ID, date and time from the system clock, priority, status, a one-line description, links to the job folder and to each deliverable. The lifecycle rules are phrased as thresholds. A job has not started until its In Progress row exists. A deliverable is not done until its Completed row exists, with links. And status or backlog questions are answered by reading those files, never from memory.

The Completed dashboard for the fictional SmartSaver test client: twelve delivered jobs covering all eleven pipelines, each row stating its scope and gate evidence with links to the final deliverable.

Call it the designated status register rather than the truth. Rows can go stale; a log can contradict a row. What keeps the register from drifting is that the row writes are written into the pipeline as steps rather than left as a separate chore. A continuation session re-reads project.md and reconciles against the job folder before it resumes. The SmartSaver run bent this once, honestly: the eleven pipelines were exercised as parallel jobs, and the orchestrating context wrote the dashboard rows rather than each job writing its own, which the conventions allow for parallel sub-agent work. The project logs say so in their friction notes. That's the behaviour I want from a status layer: when it deviates, the deviation is written down where the next reader will find it. In operating-model terms, the folders and the dashboard are the information-access and cadence layer; the triage and the gates sit on top of them.

The eleven pipelines

Eleven pipelines, and each one answers the same five questions about a different artifact: what comes in, what goes out, which packaged procedures do the work, what the gate refuses, and what the PM or BA still owns. The table is the lift-out version; the sections after it carry the detail, including what the SmartSaver run produced in each.

Pipeline Input Deliverable Gate refuses to ship unless
P1 Meetings transcript, notes, or a meeting to prepare for meeting summary, agenda, private brief, cross-meeting synthesis, stakeholder update every decision, action and quote traces to a transcript passage; no invented owners or dates
P2 Research research questions and constraints market-research report (plus market sizing on request) every load-bearing claim has a live link and access date; matrix cells sourced or marked unknown
P3 Discovery to proposal discovery transcripts, client materials, prior P1/P2 outputs brief and proposal every pain traces to a source; scope maps pains to solution elements both ways; no invented facts, quotes or budgets
P4 Requirements brief, research, meeting summaries BRS, PRD or feature brief requirements trace to inputs; no invented thresholds or IDs; open points tagged, not resolved by fiat
P5 UX discovery research, transcripts, briefs personas, empathy and journey maps, storyboards, interview synthesis assumptions and composites labelled; each persona section marked research-based or assumption-based
P6 Prototype requirements or brief, brand materials self-contained prototype and token sheet journey validation passes on the core flows; brand fidelity when a design system exists
P7 Presentation the deliverables to present HTML deck (PDF optional) every number traces to the underlying deliverable or a source
P8 Release notes changelog, commit log, ticket export customer-facing and internal notes (launch checklist optional) every item traces to a changelog entry; breaking changes and known issues surfaced
P9 Demo video prototype stills or app screenshots narrated MP4 and its script spec check passes; audio track is non-silent in each beat window; every spoken claim traces to the prototype or BRS
P10 Roadmap candidate items from a BRS, research, proposal or backlog scored roadmap and an HTML board every item sourced; inputs and scores shown per item; no committed dates; dependencies noted
P11 Measurement strategy, specs, OKR drafts, experiment or survey data OKR, hypothesis, experiment, instrumentation, dashboard or survey artifacts no fabricated baselines, targets, sample sizes or scores; confidence labels; OKR scores never tied to compensation

P1. Meetings: prep, intake, synthesis, comms

The meetings pipeline routes by where the meeting sits in time. Before it, an attendee-facing agenda or the user's private strategic brief (never shared with attendees). After one meeting, the core lane: the transcript or notes go into Input Data and the intake procedure produces a summary with a TL;DR, decisions, action items with owner and due date, open questions, risks and follow-up routing. Across many meetings, a synthesis of patterns and stalled threads. Outward, an async update for stakeholders who weren't there.

The gate is traceability in its plainest form. Every decision, action and quote traces to a transcript passage; no invented owners, dates or commitments; anything unresolved lands in Open Questions rather than being dropped. When a stakeholder update translates something technical into business language, that translation is flagged for the user to verify.

On the SmartSaver kickoff this produced fourteen decisions, six actions and three open questions, each anchored with a timestamp into the transcript. Those timestamps turned out to be the most reused objects in the whole run; the requirements, the proposal and the roadmap all point back at them.

What the PM still owns: which meetings matter enough to summarise, who gets the update, and whether a "decision" in the transcript was actually decided.

P2. Market and competitor research

Input is the research question and its constraints: region, segment, which competitors must be included. The research procedure runs its own web sweep, or delegates to a multi-source sweep with adversarial verification when the request asks for that depth, and load-bearing figures get a fact-check pass before they enter the report. Market sizing (TAM, SAM, SOM, the investment case) is a separate procedure that triangulates across frameworks and attaches a confidence label to each figure.

The deliverable is a report with a market overview, a competitor matrix, the pricing and discount picture, positioning, opportunities, a recommendation and a source list. The gate: every load-bearing claim carries a live source link and an access date, and every cell of the competitor matrix is either sourced or marked unknown, never guessed.

The SmartSaver report covered thirteen players with thirty-four sources, and its section numbers became citation anchors downstream in the same way the meeting timestamps did.

What the PM still owns: the question. A research pipeline with a traceability gate will answer precisely the question it was given, which is a reason to spend longer on the question.

P3. Discovery to proposal

This is the longest chain in the workspace, meetings into research into brief into proposal, and it ends by deciding things no transcript contains: scope, approach, timeline, team, price. So in Manual mode its first production step is the grilling interview, on what the engagement is for, where scope stops, what's explicitly out, the delivery model, and the assumptions any estimate would rest on. The model reads the inputs itself first; facts are never the user's job to recite.

The intermediate artifact is a pain-point register (pain, evidence, impact, priority), followed by a scoped competitor-solution scan, a brief, and the proposal: understanding, proposed solution, scope and deliverables, approach, team, assumptions, next steps. Diagrams are routed by type (a flow to one engine, a current-state systems view to the other) and the brand check applies the moment one is rendered.

The gate has three clauses: every pain traces to a source. Proposal scope maps pains to solution elements in both directions, so there are no orphan pains and no orphan scope. No invented client facts, quotes or budgets. On the SmartSaver run that meant ten pains mapped to nine solution elements with a two-way coverage check, verbatim quotes with their timestamps, and a budget line that reads "undisclosed" because the kickoff never stated one. The proposal ran in Auto, so everything grilling would have asked became a labelled assumption instead.

What the PM still owns: the engagement shape. The pipeline can tell you which pains the evidence supports; it can't tell you which engagement you want to sell.

P4. Requirements: BRS, PRD, feature brief

A requirements document's entire job is to state decisions, which is why every branch still unsettled when drafting starts turns into an invented threshold the gate has to catch. So P4's first step routes by what is actually missing. If the solution space is still open, a brainstorming procedure widens it, and that branch runs in Auto as well, because it's a stated default rather than an interactive extra. If the approach is agreed but boundaries, states and thresholds aren't, the grilling interview narrows it, Manual only, three rounds at most. Both true: widen first, then narrow.

Two lanes follow. The light lane produces a one-to-two-page feature brief with in and out of scope and a measurable success metric. The full lane runs requirements engineering in EARS-derived form, then either the BRS or the PRD procedure. (Canonical EARS writes an event-driven requirement as "When trigger, the system shall response"; the workspace's procedures use a WHEN/THEN/SHALL variant of it.) The document's process and flow diagrams are emitted as Mermaid source and handed to the flow-diagram engine, which relayouts them and fails mechanically on overlaps and collisions. Add-ons on request: a whole-feature failure-mode catalog, and Given/When/Then acceptance criteria when a client's QA team prefers that to EARS. If you came looking for an AI PRD generator, this is the nearest thing the workspace has. The difference is the gate: requirements trace to inputs, no invented thresholds or IDs, open points tagged rather than resolved by fiat.

The SmartSaver BRS has twenty-three acceptance criteria, and every one ends with an anchor: a timestamp into the kickoff transcript or a section number in the research report. "WHEN conducting the Adaptive Intake THEN the system SHALL ask between 3 and 5 questions. [14:04]" is a representative line. The log records an illustrative "40 percent claimed, 12 percent real" discount example being kept out of the criteria, because it was an example rather than a threshold. It records three edge cases logged as open questions, because the inputs didn't resolve them.

Acceptance criteria from the P4 requirements job in EARS-derived WHEN / THEN / SHALL form, each one ending with a timestamp anchor into the kickoff transcript it traces back to.

What the PM still owns: the scope boundary, the open questions, and the call on whether a kickoff remark was a requirement or a wish.

P5. UX discovery artifacts

Input is whatever grounding exists: research, transcripts, briefs. When the ask is open (how should these users be understood at all?), a research-methods procedure and a double-diamond framing pick the method. Otherwise the artifacts are produced directly: personas, empathy maps, journey maps, storyboards, a synthesis of user interviews across participants (a different procedure from meeting intake), and a jobs-to-be-done canvas. Deliverables are markdown or styled HTML.

The gate is a labelling rule. Artifacts are grounded in the available inputs, and every assumption or fictional composite is marked as such; a persona is tagged research-based or assumption-based section by section. The SmartSaver journey map renders that rule as a grounding key: R for a statement traced to the meeting summary or the research, A for an assumption added for narrative concreteness, with the tag on every card.

The P5 journey map delivered as HTML: an emotion curve across five stages, with a grounding key tagging every statement R for research-based or A for assumption.

What the PM still owns: whether a composite persona is good enough to design against, or whether the A tags are telling you to go and run the interviews.

P6. Prototype

A prototype is one of the most expensive things in the workspace to get wrong, because the mistake only becomes visible once it exists. Which screens, which flows, which states (empty, loading, error, success), how far the fidelity goes, real or placeholder data: none of that is determined by a brief. So in Manual mode the grilling interview runs before a single screen is built, capped at three rounds, with settled decisions written to the working folder and open ones flagged beside the prototype.

Build order starts with the design system (client tokens, or a one-off token sheet for a test client), then a design direction chosen with a taste procedure and a design-intelligence search. Then the prototype build, component-level styling, optional motion. Journey validation with Playwright runs last: drive the core flows, screenshot each state, run an accessibility and heuristic audit. The deliverable is a self-contained prototype that runs from a link, plus its token sheet.

The SmartSaver prototype cleared the gate under scripted conditions. The three-stage flow (adaptive intake, ranked offers, the savings calculator) was driven end to end in headless Chromium at a 390 by 844 phone viewport: sixteen screenshots, an in-page accessibility audit, zero console errors. Two details from its log are the kind I look for. The real retailers used as sample data are never labelled untrustworthy; the low-trust and inflated-discount examples are attached to clearly fictional seller names, so the trust features are demonstrated without a defamatory claim. And the "why this order?" sheet gives one test case consistent with the ranking rule from the BRS: the affiliate with the highest payout ranks sixth.

The first screen of the P6 clickable prototype at a phone viewport: a guided intake capped at five questions, the same cap the P4 acceptance criteria specify.

What the PM still owns: which screens exist at all. The validation gate proves the scripted flow completes end to end at a phone viewport with zero console errors, which is a different question from whether it's the right flow.

P7. Presentation

Input is the content to present: research, a BRS, status, a proposal. An artifact-planning procedure sets the structure, a slides procedure builds a self-contained HTML deck on the client's tokens when a design system exists, and a PDF export is offered. The gate is short and unforgiving: every number and claim in the deck traces to the underlying deliverable or to a source, and brand fidelity holds when a design system exists.

The SmartSaver deck is eleven slides with sourced figures. Its competitive-matrix slide is built from the research report, and each cell is marked verified, absent or partial with the source identifiers in a footnote, rather than asserting gaps the research didn't establish.

Slide four of the P7 stakeholder deck: a competitive-landscape matrix built from the P2 research report, each cell marked verified, absent or partial with a footnote citing the source identifiers.

What the PM still owns: the story. A deck whose every number traces cleanly can still tell the wrong story, and the pipeline has no opinion about which story the room needs.

P8. Release notes

Input is a changelog, a commit log, a ticket export or a feature list. The release-notes procedure produces a customer-facing document and an internal summary. For a significant cross-team launch it offers a companion launch checklist with owners, dates and go/no-go criteria, and it skips the offer for small single-team changes.

The gate: every published item traces to a changelog entry; breaking changes, security fixes and known issues are surfaced, never buried; no invented features or dates. The SmartSaver v1.0 notes were built on a twenty-row ledger, with every highlight carrying its ticket identifier.

What the PM still owns: what the release is about. The ledger is a complete record of what shipped, and choosing the three items a customer should actually notice out of twenty is a separate judgment nobody automated.

P9. Product demo video

A demo video is narrated by default. A silent slideshow of UI stills is not the deliverable, and the order is fixed: script first, then voice, then visuals timed to the voice. The script is written beat by beat, benefit-first and honest, sized to the workspace's default of about three words a second. The voiceover helper warns on any beat whose rendered audio overruns its window, so the line gets tightened. That helper also times each beat, normalises loudness and muxes the narration onto the rendered video. The default engine is the operating system's built-in voice, with ElevenLabs as an option when a key is present in the workspace. Then a fresh Remotion project is scaffolded per job, the beats are built from isolated UI slices paced to the script, and the video is rendered.

The gate is mechanical where it can be. A spec check on duration and resolution with ffprobe. A non-silent audio track, measured with ffmpeg's volume detection (mean volume well above minus 80 dB) and a silence-detection spot check that each beat window is non-silent. A script in which no claim is absent from the prototype or the BRS. UI slices never restyled. The SmartSaver video is nineteen point eight seconds of narrated, portrait video, and its script table maps each spoken beat to the screen it shows and the acceptance criterion it rests on.

What the PM still owns: what's worth twenty seconds of a prospect's attention.

P10. Roadmap and prioritization

A roadmap is nothing but judgment calls: which framework, what a horizon means here, what "done" looks like, whose priorities win a tie. None of that sits in the input folder, and the gate requires the scoring inputs to be shown per item. So in Manual the grilling interview runs before anything is scored. It asks for the scoring rules (what counts as high reach, which effort bands) rather than a score per item, which would burn the three-round budget on a twenty-item backlog.

The roadmap procedure inventories the candidates, each traced to a source, scores them (RICE by default; MoSCoW, Kano or WSJF on request), assigns outcome-based Now, Next and Later horizons, and renders a self-contained board. Timelines, Gantt charts and theme trees go to the engine that can render them. When the candidate list doesn't exist yet, an opportunity-tree procedure discovers it first; a change of direction gets a pivot-or-persevere record.

The gate: every item sourced, inputs and scores shown per item, no committed dates in Now/Next/Later, dependencies noted, no invented items or deadlines. SmartSaver's board has seventeen RICE-scored items, seven, seven and three across the horizons. Two of them carry an override badge. RICE as operationalized in this run under-scored two enablers that are hard prerequisites for the core loop. The log shows them placed by dependency with the reason stated, rather than a number inflated to make the board look right.

The delivered P10 roadmap board: seventeen RICE-scored items across Now, Next and Later, with dependency chips and an override badge where an enabler was placed by dependency rather than raw score.

What the PM still owns: the tie-breaker. The pipeline can show you that two items score the same; it can't tell you whose quarter gets the disappointment.

P11. Measurement and experimentation

P11 routes by lane, and every lane is text-only. OKRs: draft or coach a set, or grade a completed one at cycle close. Hypothesis to experiment: frame the testable hypothesis, design the test (variants, sample size, duration), then analyse and document the result. Analytics: an event-tracking contract first, a dashboard specification on top of those events second. Surveys: segmented findings with confidence labels.

The gate here is the strictest in the workspace, and the procedures enforce part of it natively. No fabricated baselines, targets, sample sizes or scores (the procedures refuse, and the instructions say to honour the refusal rather than override it). Every number traces to the inputs or a live source. Statistical claims carry confidence labels. OKR scores are never tied to compensation or individual performance.

The SmartSaver measurement job ran in Auto and produced four artifacts: a twelve-event instrumentation spec, a nine-metric dashboard requirement with the greenfield baselines left empty, a hypothesis, and a fifty-fifty experiment design. It was also the job where the finals gate did its most visible work, which is the next section.

What the PM still owns: which bet is worth measuring, and the uncomfortable admission that a baseline you don't have is a baseline you don't have.

Grilling sits in exactly four pipelines, and the other seven leave it out on purpose

The grilling interview is wired in as the first production step of four pipelines and kept out of the other seven on purpose. P3, P4, P6 and P10 decide things the inputs don't contain: the shape of an engagement, the boundary of a scope, which screens to build, which framework and which tie-breakers. Interviewing the user there, in rounds, each question carrying a recommended answer, reduces the chance of the pipeline producing something the user didn't expect. A grilling run after the work is done can only report that it went the wrong way.

The other seven, P1, P2, P5, P7, P8, P9 and P11, transform a source into an artifact and are bound by traceability gates. Interrogating the user there manufactures decisions the evidence doesn't support. That doesn't mean those pipelines carry no judgment; a persona or an experiment design involves real choices. It means the choices are exposed as labelled assumptions and open questions, or deferred in Auto, rather than elicited from the user and then passed off as sourced.

Three neighbours, three jobs. A brainstorming procedure opens the space. Grilling narrows it to settled decisions before the artifact exists. The adversarial review attacks the finished artifact afterwards. The interview has a hard ceiling I set at three rounds beyond the two or three rounds of triage, a budget chosen to keep intake short rather than a measured optimum. Hitting the ceiling with questions left is a signal to write down what's settled and carry the rest into the deliverable's open questions, not a licence to keep asking. It opens no dashboard row of its own, it runs only in Manual, and a resumed session re-reads the settled decisions rather than re-grilling from scratch.

The finals gate is adversarial, and it runs in both modes

Every client-facing final (a proposal, a BRS, a research report, deck content, release notes, a roadmap, a measurement artifact) passes an adversarial review before it lands in Final Deliverables. The review procedure, utility-pm-critic, dispatches the pm-critic critic as a separately-contexted run: same model family, fresh context, an explicit brief to attack the artifact. The fresh context removes the drafting conversation's anchoring; it does not remove the model family's shared blind spots, which is one reason the record below matters more than the mechanism. Findings come back with severities. P0 and P1 findings are resolved, or explicitly accepted by the user, before delivery, and the outcome is logged in project.md in a fixed shape: "Gate: pm-critic PASS/FAIL, N findings (P0: X, P1: Y), resolution". In Auto the review runs without asking: fix, re-review, and an unresolved P0 means stop and report.

I keep seeing the same pattern when this gate runs on quantitative artifacts, and the SmartSaver measurement job is the cleanest instance I have. Four artifacts went to review in parallel, one critic per artifact, and the findings were summed across the four: forty-eight in round one, four of them P0, nineteen P1, fifteen P2, ten P3. (The severities are the critic's own; the log states no deduplication rule, so treat the total as raw logged findings.) The four P0s were the ones a fluent draft hides best. An experiment's primary metric was defined in a way only the treatment arm could produce, which rigs the comparison; it was rewritten as a symmetric any-referrer metric. The control arm carried a legal exposure, resolved with a neutral "not verified" flag and a legal sign-off added as a launch blocker. A data-transfer gap under GDPR became a residency precondition. And a threshold of "at least four weeks" had been fabricated; it was removed and deferred to an open question. None of the four read as wrong. Each was a complete sentence in a confident document.

The project.md session log for a P11 measurement job, with the highlighted lines recording the mandatory pm-critic finals gate: forty-eight findings across four artifacts, cleared over two rounds.

Round two was a focused re-review. It confirmed the round-one P0 and P1 findings resolved. It also found that the fix pass itself had introduced nine new findings, seven of them P1: a metric floor with a chicken-and-egg dependency, a key defined two ways across documents, a guardrail that contradicted itself between artifacts. Those were resolved, a cross-artifact consistency check followed, and the artifacts were promoted with no unresolved P0 or P1. The log records no third independent round, so the honest summary is: the second round's fixes were verified by a self-check rather than by a third critic.

My first read of that record was that the drafting model had failed. The second read said otherwise. The drafts were fluent, internally consistent and confident. What this run's record shows is that fluency and correctness came apart on numbers, definitions, thresholds and legal exposure, and that a gate which can't be skipped in hands-off mode was the last control between that fluency and a client. This is the review-and-control layer of the operating model, and it's the one I would build first if I were starting over.

Source discipline is a rule about what the model may not do

The rule sits at the top of the conventions, it's written as prohibitions, and requirements traceability is its most visible application. External facts carry a live source link; internal facts trace to an artifact in Input Data; quotes are verbatim; no invented statistics, prices, names, identifiers or dates. Whatever can't be sourced is labelled an assumption or moved to open questions.

Traceability isn't a new idea. The BABOK Guide (version 3) carries a Trace Requirements task as part of its requirements life cycle management knowledge area. The EARS syntax came from Mavin and colleagues at Rolls-Royce, first published in 2009, and was designed to reduce the ambiguity that lets a requirement say less than it appears to. The P4 procedures use a WHEN/THEN/SHALL variant of its event-driven pattern. What the workspace adds is enforcement at the level of the individual acceptance criterion: each one ends with a timestamp or a section reference, and the gate rejects a document where one doesn't. That enforcement buys provenance and nothing past it. A sourced claim can still rest on a source that is stale, wrong or thin, which is why the research pipeline pairs traceability with a fact-check pass and the measurement pipeline pairs it with confidence labels. Provenance tells a reviewer where to look. It doesn't do the looking.

Sixty-one entries, routed and never browsed

The skills inventory has sixty-one entries. Two of them are a wrapper and a dependency, so call it fifty-nine working procedures. Grouped by the pipeline that invokes them: five for meetings, two for research, three for discovery, six for requirements, eight for UX, six for prototyping, three for presentations, two for release notes, four for video, three for roadmapping, eight for measurement. That is fifty. The remaining eleven are cross-cutting: the critic, the interview and its wrapper, the brand-voice check, an architecture-decision-record writer, audience-tailored briefings, a prioritized-action-plan procedure, the design-system builder and its dependency, and the two diagram engines. For readers who want the primitives defined properly, the mental model for skills, sub-agents, hooks and MCP does that. Here it's enough that a skill is a packaged procedure the model follows and a sub-agent is a separately-contexted run. The phrase AI agents for product managers usually means a grab-bag of assistants. In this workspace the agents are invoked by the pipeline that owns the deliverable, never picked from a menu.

The skill inventory grouped by the pipeline that invokes it, with a separate cross-cutting card holding the adversarial pm-critic gate, the grilling interview and the two diagram engines.

Two mechanics make that routing hold. Most of the skills were vendored from other workspaces and keep their own path conventions, so a mapping table in the instructions outranks each skill's native defaults. This one's "specs" folder is that job's Final Deliverables; that one's repo-root script runs by absolute path. And diagrams have two engines, chosen by what's being drawn rather than by preference. One engine owns five renderer types (architecture, workflow, sequence, data flow, lifecycle) and is preferred there because it fails mechanically on overlapping nodes and colliding labels. The other owns the remaining twenty-two of its twenty-seven types: timelines, swimlanes, quadrants, Gantt charts, org charts and the rest. A sanity check keeps the split honest: if the engine's type files don't count to twenty-seven, a type has gone unrouted and the list is re-derived rather than guessed. The flow engine is always invoked by absolute path, because a job's working directory is a client folder and a bare relative path fails silently, so the validation gate never runs at all. The absolute-path rule exists to close that gap.

What the human still decides

Put the eleven "still owns" lines together and a shape appears: the engagement; the scope boundary and the open questions; which screens exist; the story a deck tells; which three things a customer should notice in a release; what earns twenty seconds of a prospect's attention; whose quarter absorbs the tie-breaker; which bet is worth measuring; whether a P0 finding is accepted rather than fixed. And, underneath all of it, the point at which a clean run on a fixture stops being read as evidence of a mechanism and starts being read as evidence of results.

None of that is what the gates are for. A pipeline with a traceability gate answers the question it was given and answers it honestly, which moves the weight of the role onto asking the right question. That's the role redesign the playbook argues for, seen from the inside of the tooling.

Where AI product management breaks

Failure modes I watch for can each be read as a violation of a refusal rule, which is a useful way to remember them.

  • Running a pipeline from recollection. The step that was added last week isn't in the model's memory of the pipeline.
  • Reading Auto as "the AI decides". Auto decides defaults and labels them; it never decides facts, and it never skips a gate.
  • Storing brand anywhere global. With a single shared style guide, the first client onboarded becomes every later client's default.
  • Answering a status question from memory. The register exists because model memory isn't an authoritative status source.
  • Letting the finals gate become a checkbox. The P0s above were complete, confident sentences; a review limited to surface errors could miss defects of that kind.
  • Treating a fictional client's clean run as production evidence. It demonstrates the mechanism. That's all it demonstrates.

Key takeaways

  • Route first. One primary pipeline per request, chosen before any research or writing, with the pipeline's file read rather than remembered.
  • Auto mode transfers plan approval, not the release decision. The objective gates run in both modes and a failure stops the run.
  • Traceability is enforced per claim and per acceptance criterion, and it supplies provenance rather than correctness; fact-checks and confidence labels add further controls without guaranteeing it.
  • The adversarial finals review is where fluent drafts meet their defects, and its log line (gate, verdict, P0 and P1 counts, resolution) is the inspectable artifact.
  • Interviews belong where decisions are made (P3, P4, P6, P10) and nowhere else.

What a CTO can inspect

Almost none of this is visible in the deliverables themselves. A BRS with timestamps on its criteria looks like a slightly fussy BRS. What's inspectable is the record around it: the triage line that names the pipeline, the mode that was picked, the brand-check line, the gate line with its counts, the dashboard row with its links. If someone asks me whether the PM function's AI output can be trusted, that record is where I'd start, before any sample deck. It shows that the process was followed, and a reviewer still has to sample the deliverables and check the sources behind them.

So the next action, for anyone building a workspace like this, isn't a prompt. It's the refusal rules, written down before the first pipeline runs: what the model may not invent, which review it may not skip, and whose memory doesn't count as status.

AI Transparency Notice: This article and its accompanying images were created with the assistance of generative AI. The author directed the content, contributed the underlying ideas and analysis, and reviewed the final publication.

Frequently Asked Questions

What does an AI operating system for product managers actually consist of?

Three layers: a triage that classifies every incoming request into exactly one of eleven pipelines before any research or writing starts, two execution modes that decide who owns the taste calls, and a set of gates that run in both modes and refuse to ship unsourced output. The eleven pipelines are the standard BA and PM artifacts: meetings, market research, discovery through proposal, requirements, UX discovery, prototypes, presentations, release notes, demo videos, roadmaps, and measurement. Around them sit cross-cutting recipes that are not jobs on their own, such as diagrams, fact-checking, and the adversarial review. Mapped onto the seven operating-model components, the triage is the workflows-and-handoffs layer, the execution modes are decision rights, the gates are review and control standards, and the job folders plus dashboards are information access and cadence. The gate layer is not the whole operating model, and treating it as such is the common mistake.

Does running a pipeline in Auto mode mean the AI decides?

No. Auto transfers approval of the plan, not approval of the output. It decides the defaults and labels them as defaults; it never decides the facts, and it never skips a gate. In Auto the pipeline runs end to end on its stated defaults and presents the finished deliverable, so the model asks no mid-run questions. Structure, tone, visual direction and depth come from the pipeline's defaults instead of from the person. Source traceability, validation passes and brand fidelity are required to run exactly as they would in the interactive mode, and a gate failure means stop and report rather than ship past it. One honest caveat belongs with that: this describes what the instructions require and what the run logs show, not a mechanical guarantee. A person retains exception authority, and a procedural control is only as strong as the habit of not routing around it.

How do you stop an AI-generated PRD or BRS from inventing requirements?

Write the rule as a prohibition and enforce it at the level of the individual acceptance criterion, not the document. External facts carry a live source link, internal facts trace to a supplied input, quotes are verbatim, and anything that cannot be sourced becomes a labelled assumption or an open question. The enforcement detail is what makes it hold. Every acceptance criterion ends with an anchor, either a timestamp into the source transcript or a section number in the research report, and the gate rejects a document where one does not. That turns "no invented thresholds" from an instruction into something a reviewer can check by scanning the right-hand edge of the page. The second half is what happens when the inputs genuinely do not resolve a branch: it goes into Open Questions rather than getting settled by an invented number. In a run against a fictional test client, an illustrative discount example was deliberately kept out of the criteria precisely because it was an example rather than a threshold.

Why do only four of the eleven pipelines interview you before they start?

Because only four of them decide things the inputs do not contain. Discovery, requirements, prototypes and roadmaps settle the shape of an engagement, the boundary of a scope, which screens exist, and which framework and tie-breakers apply. None of that sits in an input folder. The other seven transform a source into an artifact and are bound by traceability gates instead. Interrogating the user there manufactures decisions the evidence does not support, which is the opposite of what those pipelines are for. That does not mean they carry no judgment, because a persona or an experiment design involves real choices. It means the choices surface as labelled assumptions and open questions rather than being elicited from the user and then presented as sourced. The interview also has a hard ceiling of three rounds, and hitting it with questions left is a signal to write down what is settled and carry the rest forward.

Can an AI PRD generator give you requirements traceability on its own?

Not on its own. Traceability gives you provenance, which tells a reviewer where to look. It does not do the looking, and it does not make a source right, current or sufficient. Requirements traceability is not a new idea. The BABOK Guide version 3 carries a Trace Requirements task inside its requirements life cycle management knowledge area, and EARS, first published in 2009 by Alistair Mavin and colleagues at Rolls-Royce, constrains requirements syntax to reduce ambiguity. Canonical EARS writes an event-driven requirement as "When trigger, the system shall response". What a generator adds is speed; what it cannot add is the surrounding control set. That is why a research pipeline pairs traceability with a fact-check pass, a measurement pipeline pairs it with confidence labels, and every client-facing final passes an adversarial review before it is delivered.

Where does AI product management break first?

At the review layer, and specifically on quantitative artifacts rather than prose. Fluent drafts fail on numbers, definitions and thresholds while reading as completely correct, which is exactly the failure a surface-level review misses. The adversarial finals review is the control for it. A separately-contexted critic run, with a fresh context and an explicit brief to attack the artifact, returns findings with severities, and the severe ones are resolved or explicitly accepted before delivery. In one measurement job against a fictional test client, that review surfaced defects including a primary metric only one experiment arm could produce, a legal exposure in the control arm, a data-transfer gap under GDPR, and a threshold that had simply been fabricated. None of the four read as wrong. Each was a complete sentence in a confident document. The second round then found that the fix pass had introduced new findings of its own, which is the argument for running the review twice rather than once.

More on AI Enablement