Skip to content
← All work

AI in the product workflow

I designed, shipped and ran four agent workflows that the whole company picked up, saving five hours of manual triage a week and cutting triage cycle time by 17%.

Building agents my team actually adopted, and being clear about where they stop.

5 hrs
of manual triage saved every week

At a glance

Role
Product Manager, and the builder, operator and first user of each agent
Team
Me building, with engineering, design and research coordinators as the customers
Timeline
Q1 2026 to today, still running
Users
PMs, engineers, designers and research coordinators across the company
Metrics
5 hours of triage saved weekly, 17% faster triage cycle, company-wide adoption

Problem and stakes

Most AI-focused PM roles ask for experience running an agentic experience for customers. My agent work so far is internal: the users are our PMs, engineers and research coordinators. I would rather say that plainly than hide it. I have been the builder, the operator and the first user of agents in production, so I have seen where trust breaks and how to design around it.

What follows is four agent products with internal customers. Each has a problem, users, a way of working, a hand-off point where a human stays in charge, and a result. Read it as one story that follows the product cycle: validate ideas, triage and build, then learn from each release.

What I built

Four agents, in the order of the product cycle. Each one names where the human stays in charge.

01

User persona bank

Validate ideas

The problem

Real user interviews are the best signal, but they take time to schedule, and early ideas often need a quick gut check first. Edge cases and missing user types were getting caught late, in QA or after launch.

Who it is for

PMs, designers and engineers who want early feedback on an idea, a flow or a PRD.

How it works

  1. 1Each persona is built from real interview notes from past users.
  2. 2A persona carries that user's core goal, their biases and how they judge a product.
  3. 3The team runs a new idea or PRD past several personas to spot gaps, edge cases and assumptions before taking it to real users.

Where the human stays in charge

It does not replace interviews or any other discovery method. It tells you what to ask in interviews, not what the answer is, so the team spends interview time on the questions that matter.

Results

  • Caught the assumption that every diary entry is made by the patient, when caregivers often fill it in for them, which added a caregiver role to the spec before real interviews.
Example persona card
"Maya", newly diagnosed caregiver
Sample persona, not built from real interviews
Core goal
Keep a record good enough to show the specialist without spending more than a minute a day on it.
Biases
Distrusts anything that looks like it exists for the research team rather than for her. Abandons a flow at the first required field she does not understand.
How she judges a product
"Did it give me something back this week?"

Sample card. Names and details are illustrative and not representative of the platform's users.

02

QA triage pipeline

Triage and build

The problem

QA intake was spread across Teams messages, Confluence pages, other Jira boards and an external Google Sheet. Engineers spent hours sorting it before any real work started, and duplicates slipped through.

Who it is for

Engineers, PMs and research coordinators. I treated them as the customers of the agent.

How it works

  1. 1Pulls reports from all four sources.
  2. 2Writes Jira tickets and removes duplicates.
  3. 3Categorizes each ticket.
  4. 4Investigates the codebase for a first pass: adds context, suggests a fix, or flags that someone else needs to be involved.
  5. 5Routes open questions to the right PM or research coordinator before engineering picks it up.

Where the human stays in charge

It is a first pass, not a replacement for engineers. Engineers review every investigation before anything is acted on.

Results

  • Saves five hours of manual triage a week.
  • Cut triage cycle time by 17%, measured from report to a ticket an engineer can start on.
  • Picked up by engineering, product, research operations and client success, processing about 20 tickets a week.
  • The routing step means PMs answer product questions before engineering time is spent.
How a report moves through the pipeline
Sources
Teams messages
Confluence pages
Other Jira boards
External Google Sheet
Agent
  1. 1Pull reports from all four sources
  2. 2Write Jira tickets, remove duplicates
  3. 3Categorize each ticket
  4. 4Investigate the codebase: context, suggested fix, or flag for help
  5. 5Route open questions to the right PM or coordinator
Human: PM or coordinator

Answers product questions before engineering time is spent.

Human: engineer

Reviews every investigation. Nothing is acted on without it.

03

Release insights dashboard

Learn from each release

The problem

We had plenty of signal about what went wrong in each release, but it was spread across the Jira board, QA documents and feedback in Confluence. Nobody had time to read it all and look for patterns, so the same pitfalls came back release after release.

Who it is for

PMs writing specs, engineering leads planning work, and leadership reviewing release quality.

How it works

  1. 1AI categorizes qualitative inputs (feedback, QA notes) and quantitative inputs (ticket counts, areas, types).
  2. 2It groups them into patterns of friction and common pitfalls.
  3. 3It became a living dashboard I update and present each release.

What it found

When we build patient-facing forms, we get about 20% more bug and QA tickets, and they cluster in conditional logic and data export.

What changed because of it

I now write PRDs that cover those risk areas up front, with user stories and success criteria for them, before engineering starts. Release by release, the dashboard tracks which areas need strengthening and which are consistently strong, and we could connect good release outcomes back to well-scoped PRDs.

Where the human stays in charge

Patterns guide the spec and the conversation. They do not make the decision.

Results

  • QA tickets on patient-facing form releases fell 25% once PRDs covered the risk areas up front.
  • Gave the team evidence that spec quality matters, which made the PRD framework an easier sell.
Release insights, redrawn with sample data
QA and bug tickets when building patient-facing forms
+20%
vs the average release
QA tickets on those releases
−25%
after risk areas went into PRDs
Where those tickets cluster (sample counts)
  • Conditional logic 14
  • Data export 11
  • Everything else 6
04

PRD orchestrator

Supporting piece

The problem

Writing a first PRD draft from a pile of client conversation notes is slow, and the structure varied from PM to PM.

Who it is for

PMs. The PM owns every decision and edit.

How it works

  1. 1Agents draft PRDs from real client conversation notes.
  2. 2They follow a framework built from our past hand-written PRDs.
  3. 3Building the framework got product and engineering in one room to agree on what helps in a spec and what is just extra words.

Where the human stays in charge

The PM owns every decision and edit. The draft is a starting point, never the spec.

Results

  • Drafted 14 PRDs so far, and a draft now needs about an hour of editing where a first draft from notes used to take a day.
  • Produced the failure story below, which is the part I would most want a hiring manager to read.
A success criterion, before and after the fix
Before: invented by the agent

"95% of researchers export a cohort within the first session."

Reads well. The product had no export at the time, and no client had asked for one.

After: grounded and traceable

"A research coordinator can build and save a cohort without asking support for help."

Source: client call, 14 March, "we keep having to email you to get the list." The agent now cites the note every criterion came from, or asks.

What went wrong

The PRD orchestrator added success criteria that made no sense for the codebase, the project or the client outcome. They read well, which is what made them dangerous. I caught it on review when a criterion said that 95% of researchers would export a cohort in their first session, and the product had no export at the time. I made three changes: I grounded the agent in project context so it drafts from what the product actually does, I added a review step that lists every success criterion with the source note it came from, and I tightened the instructions so the agent flags anything it cannot trace to a conversation. The rule I follow now is that an agent never gets to invent a requirement. If it cannot cite where a requirement came from, it asks instead of writing.

The process win

Building the PRD framework got product and engineering in one room to agree on what helps in a spec and what is just extra words. The dashboard later let me connect good release outcomes to well-scoped PRDs. That loop, spec quality to release quality, is the thing I am proudest of in this work.

Results

  • Five hours of manual triage saved every week by the QA pipeline.
  • Triage cycle time cut by 17%, measured on the QA pipeline.
  • All four agents adopted across the EPD (engineering, product and design) teams, not just by the dev team.
  • A hallucination problem caught, contained and turned into a rule the team still follows.

Principles I would bring

  • Measure the agent on task completion and repeat use, not on output volume.
  • Make the hand-off to a human clear and early.
  • Ground the agent in real context so it does not invent requirements.
  • Watch for where trust breaks and fix that first.

What I'd do next

  • Add a lightweight feedback signal to every agent output (useful, not useful, wrong) so repeat use and trust can be measured per agent instead of inferred.
  • Extend the release insights loop so a PRD can be scored against the known risk areas before it goes to engineering.

Why this matters

Any customer-facing agent serves people who want results but also want control. My agents serve teammates with the same need: do the tedious part, show your work, and hand off the decision. Those are the principles I would bring to an agent product for customers.

Next case study

Researcher Portal

I led a 0 to 1 researcher portal from kickoff to an on-time conference launch for a skeptical client, then worked with Sales to turn it into a configurable product that went to two more clients.

  • 0 to 1 build
  • Scope leadership
  • Client management
  • Sales partnership
Read the case study →
3
clients on a product built for one