← Blog

AI and Design Need a Shared Definition of Done

AI and design need shared release criteria. Diagnose conflicting definitions of done and build acceptance checks around real task completion.

A printed definition-of-done checklist sits beside a support reply draft, a policy page, and handwritten review notes on a worktable.

The feature passed QA. The model returns a plausible answer. The interface looks finished. Yet users generate a draft, hesitate and leave without using it. Your team may have a quality problem, but first check whether AI and design are working toward different definitions of done.

Engineering may call the feature complete when generation succeeds. Design may call it complete when the interaction is understandable. Product may call it complete when someone uses the result. Each team can meet its own standard while the user still cannot finish the job.

That disagreement usually stays hidden until launch. Then every abandoned output becomes another debate about model quality, copy or onboarding.

Before changing the feature, ask each function to finish this sentence independently: “This experience is done when the user can…” Compare the answers. If they describe different stopping points, you have a release contract problem.

What AI and design need to call done

A shared definition of done specifies the evidence required before an AI experience is considered releasable. It should cover the usable path, predictable failure states and the user's ability to take responsibility for the result.

Keep that separate from task-specific acceptance criteria. The Scrum Guide treats the Definition of Done as a quality standard for the increment. Individual workflows still need their own concrete checks.

For a support-reply feature, the general standard might require visible draft status, tested recovery paths and instrumentation through application. The specific criteria would address policy references, editing and sending.

AI and design share a useful completion standard only when it reaches beyond “an answer appeared” to “the user could safely complete the intended task.”

That does not mean every output must be perfect. It means the product handles variation deliberately. A draft can be useful but incomplete. Missing context can require clarification. An unsupported request can require a refusal. Those are valid product states, not exceptions to hide behind a generic error message.

The release question is whether those states help users proceed, correct or stop with a clear reason.

Diagnose the disagreement before changing the feature

Look at where each function stops evaluating the experience. The mismatch often explains why a team keeps improving the wrong thing.

Observable symptom Conflicting completion standards Evidence to inspect
Successful generations, few applied outputs Engineering measures response success; product needs task completion Sessions from generation through application or abandonment
Positive demo feedback, repeated verification elsewhere Design evaluates clarity; users need grounds for judgment Which claims users check outside the product
Frequent regeneration, little progress The team treats another output as recovery Whether users can correct a specific error without restarting
Strong first use, weak repeat use Launch criteria cover novelty, not recurring workflow fit Use across subsequent eligible task opportunities

These are diagnostic hypotheses, not conclusions. Watch representative sessions and ask users what prevented the next action. A low application rate could reflect poor output, missing permissions or a task that did not need AI.

When AI and design disagree about the stopping point, do not resolve it by averaging their metrics. Agree on the user outcome first, then assign supporting checks to each function.

The distinction between output generated and output applied matters here. A successful response is evidence of system behavior, not proof that the job is complete.

Write a contract for one real task

Start with a narrow workflow. “Help support agents” is too broad to test. “Draft a reply using the current ticket and approved refund policy, then let the agent review and send it” has a clear boundary.

For that hypothetical feature, a release contract could include these checks:

  • Context: The agent can see which ticket and policy version informed the draft. Missing required context produces a clarification state, not an invented answer.
  • Judgment: Policy-dependent claims have inspectable references. The interface distinguishes a suggested reply from an approved decision.
  • Correction: The agent can change one sentence without losing edits elsewhere. Regeneration does not silently overwrite reviewed text.
  • Application: Nothing sends automatically. The agent can review, edit and deliberately send through the existing support workflow.
  • Failure: Conflicting policy information and unsupported promises have tested handling. The agent has a usable manual path when drafting cannot proceed.

These are proposed criteria for this example, not universal requirements. A low-risk brainstorming tool needs a different contract from a feature affecting customer commitments.

For AI and design, the shared requirement is that each criterion has observable evidence. “Feels trustworthy” is not a release check. “Participants can locate the policy supporting the refund statement” is testable, although finding a reference alone does not prove the statement is correct.

Choose the pass conditions before testing. Otherwise, a disappointing result invites the team to redefine success after seeing it.

A printed definition-of-done checklist sits beside a support reply draft, a policy page, and handwritten review notes on a worktable.

Assign ownership to evidence, not just deliverables

A shared contract does not mean everyone owns everything. Give each check an accountable owner and name the evidence they must bring to the release review.

Engineering can own preserved edits, permission enforcement and failure-state behavior. Design can own whether users distinguish draft from approved output and recover without losing work. Product can own the task boundary, risk tolerance and outcome measures.

Domain reviewers should assess claims that require specialist judgment. Designers should not be asked to certify refund eligibility, legal accuracy or medical appropriateness simply because those claims appear in an interface.

The review should compare evidence against the agreed contract. If users cannot identify an unsupported claim, a technically successful response does not cancel that failure.

AI and design need a clear escalation rule: who decides whether an unresolved failure blocks release, narrows the supported workflow or requires another safeguard?

Record that decision. Otherwise, the same disagreement returns during the next model change or interface revision.

Test the contract across variation

One polished demo cannot establish release readiness. Use scenarios that exercise the supported task and its boundaries: sufficient context, missing context, contradictory information and a request outside the feature's scope.

Separate deterministic checks from judgment checks. Whether regeneration preserves an edit can be tested directly. Whether a reply is suitable for a particular customer requires a defined rubric and relevant reviewers.

Measure the whole attempt. Include time spent checking and correcting, not just generation latency. Compare against the existing workflow before deciding that the feature saves effort.

For AI and design, repeat-use evidence should also be tied to eligible opportunities. A user with no relevant task this week is different from one who had the task and avoided the feature.

Before release, collect evidence that the task can be completed. After release, monitor whether it actually is completed. Re-run the relevant checks when changes could affect the contract, including changes to defaults, available context or generation behavior.

Frequently asked questions

Is a shared definition of done the same as an acceptance checklist? No. The definition establishes the recurring quality standard. A task-specific checklist translates it into checks for one workflow, such as preserving reviewed text or exposing the policy behind a claim.

Does this require a retention target before launch? No. Before launch, test completion and recovery. After launch, measure application and repeat use across eligible opportunities. Define the instrumentation before shipping so those outcomes are observable.

Make the next release decision explicit

Pick one abandoned-output workflow. Write its completion contract, assign evidence owners and review the last release against it. Decide which unresolved failure blocks the next release.

If the underlying symptom is still unclear, use the free AI adoption Triage tool. For a deeper working structure, the AI Product Adoption Deck is a 104-card, 124-page diagnostic playbook with 12 diagnostics, 80 action cards across 10 stacks and 12 workshops with fillable deliverable templates.

Start with one shared stopping point. Make the next release prove it.


← All postsGet the Deck →