XY

AI workflows · Durable execution

Human Approval in Durable AI Workflows

A human-in-the-loop button is not a safety boundary if the process forgets the wait, the approval has no identity, or the model can call publish by another route.

Decision map Persist before waiting; authorize before publishing

Research and evaluation may retry. The approval hook persists the pause, and only a recorded decision can cross the side-effect boundary.

  1. 01 Research

    Parallel, retryable evidence lanes

  2. 02 Draft / evaluate

    Bounded revision loop

  3. 03 Persisted hook

    Workflow suspends without an open request

  4. 04 Human decision

    Approve, reject, expire, or cancel

  5. 05 Publish

    Side effect runs only after approval

Waiting is a state

Many AI demos keep a request open while a model works, then show an approval button in the same browser tab. Close the tab, restart the server, or deploy a new version and the apparent workflow disappears. That is a UI sequence, not durable coordination.

Durable Brief treats waiting as persisted workflow state. Research can finish, a draft can pass evaluation, and the run can pause on an approval hook without keeping one Node process alive. Refreshing the page does not authorize publication and does not erase the pending decision.

Human approval matters only when the system can remember what is waiting, who may approve it, and exactly what will happen next.

Give each phase different execution semantics

The workflow fans out three research tasks in parallel because they spend most of their time waiting on IO. It fans in before drafting because one artifact needs one coherent state. An evaluator-optimizer loop can revise the draft, but it is capped at three passes so the system cannot spend indefinitely in self-critique.

LLM calls and publication are retryable steps. The orchestration function records the sequence, while step functions define work that may be retried. That boundary matters: replaying orchestration should not silently duplicate an irreversible side effect.

research A ─┐
research B ─┼→ draft → evaluate ↺ revise → HUMAN GATE → publish
research C ─┘

The gate must sit outside model authority

A prompt that says ask a human before publishing is not an enforcement mechanism. The model may misunderstand the instruction, a later refactor may bypass it, or another tool path may expose the same side effect. In Durable Brief, publish follows a workflow hook that the model cannot satisfy by generating text.

This is the same principle as keeping authorization out of a client-supplied tenant header. The actor proposing the action does not define the authority to execute it. The workflow runtime and approval endpoint own that decision.

  • The artifact under review needs a stable version or digest.
  • The approver needs authenticated identity and an allowed role.
  • Approve, reject, edit, timeout, cancel, and escalate need explicit states.
  • Publication must consume the approved artifact—not regenerate it after approval.

A demo can be deterministic without being fake

The public demo works without a model API key. Its research and draft content are deterministic, and the evaluator is forced through a revise-then-pass sequence. That makes the durable behavior inspectable on every visit while the approval hook remains real.

This split is useful for portfolio software and production tests. External intelligence may be unavailable, expensive, or variable; the state machine should still be demonstrable. Deterministic fixtures let tests prove fan-out, revision, pause, resume, and publication order without claiming that fixed text represents model quality.

Approval creates operational obligations

Once a system can wait for a person, it needs a policy for silence. Does the request expire? Who is paged? Can another reviewer take over? Can the requester cancel it? What happens if the underlying policy changes while an artifact waits? A queue of permanent pending approvals is another form of failure.

The approval record should answer who approved which artifact under which policy and when. High-risk actions may require two reviewers or separation of duties. Lower-risk actions may use time-boxed delegated authority. Human-in-the-loop is not one pattern; it is an authorization workflow with latency and accountability tradeoffs.

What I would improve next

The current project proves durable waiting and ordering. I would next authenticate approvers, sign the decision together with the artifact version, and persist a queryable audit trail. I would add timeout, escalation, cancellation, and optional multi-reviewer quorum.

I would also version prompts and policies and enforce per-run model, tool, latency, and cost budgets. The most important improvement is conceptual: approval should not be a decorative pause near the end of an AI flow. It should be a durable, inspectable transfer of authority.

Source notes

Claims you can inspect.

  1. Code brief.ts · durable orchestration and approval ↗

    Parallel research, bounded critique, a persisted hook, rejection, and post-approval publish in the implemented workflow.

  2. Code approve route · resume boundary ↗

    The HTTP boundary that validates a decision and resumes the waiting hook.

  3. Reference Vercel · durable human-in-the-loop agents ↗

    Primary guidance on pausing, persisting state, resuming from a human decision, and retrying steps independently.

  4. Reference Vercel Academy · record and approve the decision ↗

    A production-oriented treatment of verified reviewer identity, durable waits, and idempotent approval handling.