← Decision helpers

Agency Design Canvas

Should this be a workflow or an AI agent?

Decide whether a task needs ordinary automation, AI within a workflow, or an agent—and how much freedom that agent should have.

What does this framework help you decide?

Choose the simplest approach that can do the job effectively. Answer a few questions about the work, the value of independent decisions, the consequences of mistakes and the controls you have in place.

You'll get a suggested approach, a level of autonomy, human-oversight guidance and an explanation of the trade-offs. It's a decision aid, not permission to deploy.

Describe the work

Move the sliders to reflect the actual task, not the technology you want to deploy. All questions use a 0–4 scale.

1 · Is agency actually needed?

Higher values increase the potential value of independent planning and action.

2 · Is an AI model needed?

Variable or unstructured inputs can justify AI even when an agent is unnecessary.

3 · How much authority is safe to delegate?

These questions constrain recommended autonomy; they do not create a need for an agent.

4 · What additional value does autonomy deliver?

Assess incremental benefit over the best feasible workflow, not the value of the task overall.

5 · How mature are the governing controls?

Score controls that are implemented and evidenced, rather than planned capabilities.

6 · Human oversight preference

Choose an operating pattern to test against the assessment. This preference never overrides required approvals or foundational controls.

7 · Non-negotiable controls

Shared vocabulary

Precise terms help avoid treating uncertainty as though it were the same thing as autonomy.

DeterministicGiven the same input and complete system state, produces the same output.
Non-deterministicThe same apparent inputs may lead to different outcomes; causes may include randomness, concurrency or hidden state.
ProbabilisticRepresents or reasons about uncertainty using probabilities. A probabilistic model need not sample randomly at inference.
StochasticIncludes a random process or random sampling. Probabilistic and stochastic overlap; they are not sequential autonomy levels.
AgencyAbility to select and take actions in pursuit of an objective.
AutonomyDegree of discretion and independence in planning and acting without human approval.
Workflow automationOrchestration along predefined steps, conditions and permissions.
Constrained agentGoal-directed system with explicitly bounded tools, permissions, scope and checkpoints.
Required autonomyDegree of independent action genuinely needed to deliver task value.
Permissible autonomyDegree of independence allowed by risk appetite and evidenced controls.
Justified autonomyAutonomy whose incremental value warrants its cost and residual risk.
Risk-adjusted autonomyDelegation bounded by task impact, reversibility, scope, control maturity and benefits.
Scoring and decision rules (fully transparent)

Agency need is the average of four task questions; AI need averages two interpretation questions. Delegation concern averages impact, inverted reversibility and scope. Incremental value averages two value questions; evidenced control maturity averages three maturity questions. The indicative permissible-autonomy ceiling is min(4, 4 − delegation concern + (control maturity − 2) × 0.75), floored at 0; mandatory approval caps the ceiling at 2, and missing foundational controls caps it at 1. This is a discussion aid, not a quantitative risk model or approval. Low agency need (<1.6) favours workflow or AI-assisted workflow. Higher agency need can justify an agent only if incremental autonomy value is at least 1.8 and permissible autonomy is at least 1.6; otherwise favour a workflow or model-assisted workflow until controls improve. High autonomy requires agency need ≥2.8, incremental value ≥2.8, and a permissible-autonomy ceiling ≥2.8. Where need exceeds the ceiling, the output explicitly flags an autonomy–assurance gap. Thresholds are illustrative, not validated risk scoring or deployment authorization.

Human oversight: a separate design dimension

The four architectural options do not prescribe a single human-oversight pattern. Separate who plans from who authorises actions, monitors execution and owns the outcome.

Human-in-the-loop (HITL)A person must review or authorise a defined decision or action before it takes effect. May be required at every action or only at consequential checkpoints.
Human-on-the-loop (HOTL)System acts within delegated boundaries; a human supervises and can intervene, stop or override its operation.
Human-out-of-the-loop (HOOTL)No human approval or real-time supervision is needed for each execution. The system still requires accountable ownership, limits, logs and periodic review.
Human-before/after-the-loopPeople set goals, rules and permissions before execution and inspect outcomes or exceptions afterwards. These are useful complements, not substitutes for required pre-action approval.
ArchitectureIllustrative oversight pattern
Workflow automationOften unattended execution plus exception handling. May include a mandatory human approval step.
AI-assisted workflowHuman review of uncertain AI outputs or high-impact downstream actions; routine, low-risk inference may be automatic.
Constrained agentAgent plans within limits; humans authorise specified actions, exceptions, or changes of scope. Monitoring and stop mechanisms support oversight between gates.
Higher-autonomy agentBounded unattended execution, live monitoring, escalation triggers, interruptibility and post-action audit; approval remains mandatory wherever policy requires it.

Critical distinction: human review is meaningful only if reviewers have enough information, time, competence and authority to challenge, change, or halt the action. Adding a checkbox or approval button is not, by itself, a risk control.

Behind the model: evidence, misuse and sources

An evidence-informed architecture conversation, not a scientifically validated agent-selection test. Expand any section to inspect the reasoning and follow the original references.

The science behind it Research lineage

1. Levels of automation are a design choice

Parasuraman, Sheridan and Wickens (2000) distinguish information acquisition, information analysis, decision/action selection and action implementation. They argue that each function can receive a different level of automation and should be evaluated for human performance consequences. This supports asking what an AI system may decide or execute, rather than applying a single “agentic / not agentic” label to the whole system. Original research ↗

2. Agency and uncertainty are separate dimensions

Deterministic behaviour concerns repeatability; probabilistic models express uncertainty; stochastic processes involve randomness. None, by itself, tells us who can select actions or approve execution. The distinction in this canvas is a conceptual architecture model, not a taxonomy claimed to be prescribed by NIST or ISO.

3. Governance changes permissible delegation

The NIST AI Risk Management Framework organises risk work into Govern, Map, Measure and Manage, with activities throughout the system lifecycle. Its Generative AI Profile extends that work to generative-AI-specific risks. ISO/IEC 42001 specifies requirements for an organisational AI management system. These sources support the principle that deployment authority should depend on context, risk assessment, evidence and continuing control operation—not model capability alone. NIST AI RMF ↗ · NIST GenAI Profile ↗ · ISO/IEC 42001 ↗

4. More actions introduce agent-specific failure modes

OWASP's Top 10 for Agentic Applications (2026) documents threats associated with agents that plan and act through tools. NIST has also studied agent hijacking through indirect prompt injection. These are reasons to evaluate permissions, tool boundaries, monitoring, approvals and recovery separately from the quality of an agent's prose or predictions. OWASP ↗ · NIST agent hijacking ↗

5. Human approval is not the same as effective human oversight

Human-factors research documents automation bias: people may over-rely on automated advice and fail to notice errors. A nominal human-in-the-loop checkbox cannot substitute for meaningful review time, information, authority and the ability to interrupt or reverse an action. Automation-bias review ↗ · Verification-complexity review ↗

What this instrument is — and is not Methodology

This is a transparent, qualitative decision heuristic. Its sliders operationalise four architectural distinctions: predictability, uncertainty/AI need, agency need and risk-adjusted autonomy. It compares required autonomy with an illustrative permissible-autonomy ceiling, and asks whether incremental value makes the additional delegation worth discussing.

Evidence boundary: the underlying principles are supported by research and standards; this canvas's particular questions, weights, thresholds, formula and four recommendation labels are original design choices. They have not been calibrated, independently validated, certified, or demonstrated to predict incidents or financial outcomes. A “3 / 4” is an ordinal discussion input, not a 75% probability, a risk rating or a control-effectiveness measurement.

Changing an answer can change the suggested architecture. Treat that as sensitivity analysis rather than precision. For consequential cases, compare multiple candidate architectures, collect evidence, document assumptions and require the organisation's normal architecture, risk and change approvals.

See “Scoring and decision rules (fully transparent)” above for the actual formula and cut-offs used by this version.

How this can be misused Failure modes
  • Score laundering: presenting a heuristic recommendation as “science says we must deploy an agent.” Countermeasure: record alternatives, assumptions, dissent and the human decision owner.
  • Slider gaming: overstating task complexity or control maturity to get a preferred result. Countermeasure: attach evidence, use a cross-functional review and preserve the original assessment.
  • Equating controls on paper with operating controls: a policy or checkbox does not establish tested containment, identity, recovery or monitoring. Countermeasure: require demonstrations, evaluations and named operational owners.
  • Decorative human approval: the reviewer cannot realistically verify the output, challenge it or stop execution. Countermeasure: design actionable review, appropriate information, time and override capability.
  • Conflating inference with agency: a probabilistic classifier is not necessarily an agent, and constraining an agent's tools does not make its model deterministic. Countermeasure: map inference, planning, decision rights and action authority separately.
  • False comfort from deterministic workflows: predefined steps can still implement flawed rules, act on erroneous AI output or have excessive permissions. Countermeasure: assess actual failure impact for every option.
  • Hiding residual risk: a high control-maturity slider must not cancel a disallowed action, legal duty or unacceptable consequence. Countermeasure: use independent hard gates and organisational risk acceptance; never allow a high average score to override them.
  • Missing value and opportunity cost: choosing the least autonomous option by default can prevent the system from solving the problem. Countermeasure: compare measurable outcomes of workflow, AI-assisted workflow and bounded agent trials.
External sources Follow the evidence
  1. Parasuraman, R., Sheridan, T. B. & Wickens, C. D. (2000). A model for types and levels of human interaction with automation. IEEE Transactions on Systems, Man, and Cybernetics A, 30(3), 286–297. DOI ↗ — framework for levels and functions of automation.
  2. NIST (2023). Artificial Intelligence Risk Management Framework 1.0. Official publication ↗ — govern, map, measure and manage.
  3. NIST (2024). Generative AI Profile, NIST AI 600-1. Official publication ↗ — generative-AI risk considerations.
  4. ISO/IEC (2023). ISO/IEC 42001: Artificial intelligence management system. Official standard page ↗ — organisational management-system requirements (full standard may require purchase).
  5. OWASP (2026). Top 10 for Agentic Applications. Official project ↗ — agent-specific security risks and guidance.
  6. NIST (2025). Strengthening AI Agent Hijacking Evaluations. Official research blog ↗ — indirect prompt injection and agent action risk.
  7. Goddard, K., Roudsari, A. & Wyatt, J. C. (2012). Automation bias: a systematic review of frequency, effect mediators, and mitigators. Open access ↗ — human reliance on automated advice.
  8. Lyell, D. & Coiera, E. (2017). Automation bias and verification complexity: a systematic review. Open access ↗ — challenges of verifying automated outputs.
  9. Decision-Method Selector — the companion helper for choosing how a decision should be made; shares this site's transparent “science / misuse / sources” disclosure.