The gap finder found the hole I dug

The backlog could create and complete work, but missed the lifecycle in between. A story map gap analysis exposed the trust and platform gaps.

Todoista Waiting screen listing commitments grouped by the person who owns the next action
Table of contents
  1. The pattern: create and complete, but nothing in between
  2. Correctability is a trust requirement for AI products
  3. The missing terminal state became four distinct outcomes
  4. Refinement is a decision-forcing machine
  5. Can an unverified account capture commitments?
  6. Do we store pasted meeting notes?
  7. What happens when the extractor is unavailable?
  8. A duplicate did not need deletion; it needed a sharper meaning
  9. The gap no structural audit could find
  10. Each skill is tuned to a different blind spot
  11. Related articles

Create and complete, but nothing in between.

This is part three of a three-part series about building a product backlog with Claude, five custom story mapping skills, and the StoriesOnBoard MCP server.

At the end of part two, the board looked finished: nine goals, 26 steps, 82 stories, release slices, and persona assignments.

“Looked finished” is the important part.

I ran the story-map-gap-finder almost as a formality. Instead, the story map gap analysis exposed a pattern I now look for in every AI-generated backlog.

The pattern: create and complete, but nothing in between

The report could be summarized in one sentence: you can create an item and complete it, but you can barely manage its life in between.

AI gap finder reports missing edit, delete, search, people, cold-start, and privacy stories
Illustrative reconstruction: the gap finder exposed the missing lifecycle between creating and completing an item.

The missing stories were not exotic edge cases:

  • No edit or delete after confirmation. The detected inbox allowed correction before tracking, but an accepted commitment became effectively permanent.
  • No search. After several months and hundreds of items, there was no way to answer “What did I promise Anna about the contract?”
  • No people directory. You could add a person and see their open items, but a person disappeared from navigation as soon as every item was complete.
  • No cold-start path. On the first capture, the people directory was empty, so extraction was guaranteed not to match anyone at the exact moment the user judged whether the product worked.
  • No way to delete a third party’s personal data. That left a GDPR-shaped hole in a product that stores non-user names and email addresses.

None of these is a “nice to have.” A real user would encounter most of them during the first week.

Todoista Waiting screen listing commitments grouped by the person who owns the next action
The finished Waiting view shows who owns each next action and what the manager is waiting for.

Correctability is a trust requirement for AI products

The missing edit flow looked like ordinary backlog hygiene until I considered the product’s core interaction. A model extracts a person, commitment, owner, and due date from speech or meeting notes. Sometimes it will be wrong.

If users cannot correct the extraction after accepting it, every error becomes permanent evidence that the system cannot be trusted. Eventually they stop relying on any of its outputs, including the correct ones.

For an AI-assisted workflow, edit is not merely a convenience. Correctability is part of the trust model.

This is why lifecycle analysis matters. A flat feature list may contain “AI capture” and “mark complete” and look comprehensive. A story map makes the unhandled states between those two moments visible.

The missing terminal state became four distinct outcomes

One recommendation from the gap finder was “cancel a waiting item.” At that point, a waiting item could end only when the requested result arrived or when the manager took the work back.

But there is a common third case: it is no longer needed.

That prompted a deeper decision. The product needed four terminal actions, each with a different effect on the person’s history:

ActionWhere the item goesWhat the person’s history records
ReceivedBack to My TasksShows a successful delivery
Taken backBack to My TasksDoes not count as the person’s delivery
CancelledClosedThe request was real, but the manager withdrew it
DeletedDisappearsNo history, as if it never existed

The distinction is not semantic decoration. If I delete a quotation request after cancelling a project, Anna’s history makes it look as though I never asked. If I cancel it, the product preserves the fact that the request existed while avoiding any suggestion that Anna failed to deliver.

In a product built around delegated work, the person’s history is part of the core value. These four rows became acceptance criteria on the relevant cards.

Refinement is a decision-forcing machine

Next, the story-refiner expanded selected titles into a user-story narrative and testable acceptance criteria. That exposed product decisions the specification had never needed to state.

Can an unverified account capture commitments?

Yes. It receives no reminder email, but capture still works. Blocking the core action until verification risks losing a first-time user before they experience the product.

Do we store pasted meeting notes?

No. The notes are source material, not product content. Only commitments the manager confirms are stored as structured items.

What happens when the extractor is unavailable?

The raw sentence is saved for later processing. This appears to conflict with “confirm before save,” but a stronger rule wins: capture must never fail. If the product loses a commitment once, users will stop trusting it during meetings.

These decisions were not missing paragraphs in the specification. They did not become visible until somebody tried to write criteria a tester could answer with yes or no.

A duplicate did not need deletion; it needed a sharper meaning

A later review found four notification stories under two different goals. Two were almost identical: set notification preferences and set reminder defaults.

The cause was traceable. The foundations checklist had generated a settings story while the product specification generated a reminder story. Two valid inputs had described the same area with different words.

Because the MCP toolset did not expose card deletion, we clarified the scope instead:

  • What: which email notifications should I receive?
  • When: at what time should they be delivered?
  • How much: what is the default follow-up delay?

The “Get reminded” step became the single home for notification behavior. The MVP held delivery; a later release held the configuration block.

The useful move was not inventing three features to justify three cards. It was separating three product decisions that had been bundled under the vague word “settings.”

The gap no structural audit could find

During the seventh review round, I casually wrote:

“This is a web app. There will be no push notifications for now, but email notifications are possible.”

That was the first time the platform had been stated.

Five stories already assumed a native mobile app: push notifications, phone contact import, swipe actions, a “move around quickly on mobile” step, and the mobile-first line in the Manager persona.

The push story became email. The others had to be rewritten.

The gap finder could not have caught this. It audits structure: missing journey phases, uncovered steps, lifecycle gaps, and edge cases. An unstated platform assumption is internally consistent on the board. It describes the wrong world without leaving a structural hole.

Discovery is what should catch it. A fundamentals interview asks where and how the product is used. I had skipped that interview because I already had a “complete” specification.

Each skill is tuned to a different blind spot

The full experiment changed how I think about assistant skills:

  • Discovery challenges category defaults, unclear users, and unstated differentiation.
  • The foundations pass adds privacy, errors, empty states, and other requirements that specs routinely omit.
  • Gap analysis catches missing lifecycle states, navigation paths, and correction flows.
  • Refinement forces ambiguous titles to become testable product behavior.
  • Navigation and read-back test release coherence and compare the live board with the promised plan.

The skills are valuable because they are fast, but speed is the least interesting property. Each one is tuned to find a different kind of plausible-looking mistake.

And they share the same final habit: read the board back, compare it with what you intended to create, and report where reality differs from the plan.

The five skills used in this series were product-discovery, story-map-creator, story-map-gap-finder, story-refiner, and story-map-navigator. The writing skills worked directly on StoriesOnBoard through MCP.