Skip to content

Framework

AI Agent Governance Maturity Model

Describe what a defined set of agent workflows can demonstrate, keep unverified paths visible, and choose the next result to establish.

The KonaSense AI Agent Governance Maturity Model is a six-level model for describing what an organization can demonstrate about governing a defined set of agent workflows and choosing the next verifiable improvement. It assesses the evidence supporting a governance claim, not the number of tools purchased or policies written.

The levels are Unmanaged, Discovery, Observability, Policy, Runtime Enforcement, and Adaptive Governance. They are KonaSense's proposed assessment structure, not an industry standard, certification, or empirically validated scoring system. A level applies only to the scope and conditions supported by the review.

Define what is being assessed

Name the agent workflows, execution paths, environments, and operations included in the assessment. Identify the relevant users, execution identities, configurations, policy versions, and review period. An interactive client and a scheduled job can use the same agent software while exposing different authority and controls.

Describe the required governance outcome. For an agent that changes an internal escalation roster, the outcome might require changes to remain within an approved team's records and respect a designated owner's approval. That is an illustrative requirement, not an account of a KonaSense integration. The assessment must establish what could change, whose authority applies, and which paths can produce the effect.

AI agent governance defines the broader delegation and responsibilities. The KonaSense Runtime AI Governance Framework maps the functions needed to support a governance outcome. This maturity model adds a different question: what can the scoped program demonstrate now, and what evidence would justify the next claim?

Do not average different surfaces into one enterprise score. A runtime control demonstrated on one path can coexist with an unexamined scheduled executor. Preserve the narrower finding and the wider uncertainty separately. If the scope cannot be established, report the assessment as incomplete rather than assigning a precise-looking level.

The six levels

Levels 1 through 5 build on the relevant supporting conditions below them within the assessed scope. A policy document alone does not establish Level 3 if the team cannot identify the activity it governs or obtain the observations needed to assess it. Record a later-level capability as a specific finding when the supporting conditions remain incomplete.

Level 0: Unmanaged

Agent activity lacks consistent inventory, policy, or evidence practices within the reviewed scope. Governance depends on individual knowledge and ad hoc decisions that the organization cannot reliably reconstruct or maintain.

Use this description when a bounded review establishes that those practices are not operating consistently. Relevant evidence can include the actual workflow and configuration, available records, and confirmation of responsibilities with the people involved. Failure to obtain records or speak to an owner is insufficient by itself: an environment that has not been examined is unassessed, not automatically unmanaged.

Preserve any individual rules or controls that the review does find. The first improvement is to establish an accountable inventory and scope for material agent use, with known omissions identified. It is not to replace uncertainty with a Level 0 label.

Level 1: Discovery

The organization can identify material agent usage and surfaces within the assessment boundary. The inventory relates known agents and execution paths to their owners and relevant environments, rather than listing only product names.

Support this claim with inventory sources, the method used to reconcile them, the period or configuration they describe, and the areas those sources cannot cover. Discovery by an application catalog, endpoint observation, or an operator's report may establish different facts; record which source supports which entry.

The next result to demonstrate is a coherent view of relevant activity on those paths. A discovered agent is not necessarily monitored, and an inventory entry does not establish that its tools, sessions, or actions are observable. Keep that limit attached to the Level 1 finding.

Level 2: Observability

The program can examine relevant sessions, tools, actions, and risk signals for the scoped workflows, preserving the origin and meaning of its observations. Visibility must be sufficient for the review question, not a promise to capture every event or an agent's internal reasoning.

Show that the available records can be related to the activity under review. Identify what the producer observed, what collection preserved, and where information is missing. A proposed request, an execution attempt, and a destination result remain different stages. A large volume of logs does not establish that these relationships are available.

The AI agent audit trail guide describes the evidence relationships in detail. The next improvement is to make acceptable and unacceptable behavior explicit enough to evaluate against that context. Observing a risky operation does not itself show that a governing rule exists or that the operation can be constrained.

Level 3: Policy

Rules define acceptable and unacceptable behavior for the scoped activity. They identify relevant conditions, decision authority, and the treatment of exceptions. Reviewers can determine which rule applies to a concrete use or proposed action, instead of inferring permission from the agent's general purpose.

Evidence includes the applicable policy version, its owner, the facts its conditions require, and how evaluators interpret its responses. Use concrete cases to establish whether a rule distinguishes the intended behavior from behavior outside the delegation. Evaluation can involve people, software, or both; a document named “AI policy” does not establish that its conditions are usable.

This level does not establish runtime enforcement. The next result is evidence that a specified decision can affect the protected operation at an available control point. A policy saying “require approval” and a component displaying a confirmation prompt do not, by themselves, establish that an authorized approval governs execution.

Level 4: Runtime Enforcement

Policy can affect agent actions during execution at identified control points. The claim must name the protected operation, the paths passing through the control, and the conditions under which the intended response can be applied.

Evidence must connect the evaluated proposal and decision to the executor's behavior, with destination observations where the claimed effect requires them. Examine permitted and restricted cases and relevant failure behavior. A decision record alone cannot establish that the executor withheld an operation. A control that continues when a required authorization is unavailable cannot be credited with enforcing that authorization on the affected path.

NIST MEASURE 2.3 calls for documented measurements under conditions similar to deployment and attention to differences between the measurement and deployment settings. This supports specifying the conditions behind a control claim; it does not validate this maturity level. The runtime governance pillar explains the decision and application contracts to examine.

The next improvement is maintaining that capability through operation and change, including approval handling, changing context, and evidence gaps. One successful demonstration cannot establish that the process remains effective across a review period or a different configuration.

Level 5: Adaptive Governance

Automated policy evaluation, applicable approval processes, risk context, and evidence operate as a maintained process across the AI uses and agents in scope. Changes and feedback lead to accountable review and verified adjustments, with continuing visibility into where the governance process has degraded or remains incomplete.

Adaptive does not mean that policy rewrites or approves itself. People retain responsibility for policy changes and exceptions. Automation can support evaluation, routing, context updates, and evidence collection while human decisions remain necessary where the policy requires them. Continuous operation means an ongoing process with defined coverage and response arrangements, not a guarantee of uninterrupted control over every event.

Support this claim with evidence over an appropriate review period: operating records, handled exceptions, detected gaps, authorized changes, and verification of the resulting configuration. The period and observations must fit the workflow and the conclusion. The model does not prescribe a universal number of days or treat a configured monitoring rule as proof that its response process operates.

NIST GOVERN 1.5 addresses planned monitoring and periodic review with assigned responsibilities. MANAGE 4.2 addresses continual improvement, including changes to organizational procedures. These principles support maintained governance; they do not make self-modifying policy a maturity requirement.

Build a finding from evidence, not a feature label

For each proposed level, keep a review record that identifies:

  • Scope and claim: the workflows, paths, operations, configuration, and period the finding describes.
  • Supporting evidence: its source, what was examined, and which requirement it establishes.
  • Limits and contrary evidence: exclusions, gaps, failures, or conditions that prevent a broader conclusion.
  • Decision and owner: the claim the reviewer accepts, who is accountable for it, and any operating restrictions.
  • Next result: the specific capability or evidence relationship to establish before reconsidering the level.

NIST MEASURE 2.1 emphasizes documenting test sets, methods, metrics, and tools used in evaluation. Preserve enough of that context for another reviewer to understand a control assessment. The maturity label must not replace its supporting record.

Distinguish a required mechanism that is absent from one whose behavior is unverified. If records are inaccessible, say so and identify what would permit review. If a requirement is considered not applicable, record the scope-based reason and the person authorized to make that determination. None of these states should silently become a passed requirement or an assertion that the whole environment is unmanaged.

When evidence is mixed, report a profile of supported capabilities and unresolved requirements instead of forcing one number. A documented policy can remain a valid finding even when runtime enforcement is unverified. Conversely, a successful bounded control check does not establish inventory completeness, usable policy, or a maintained review process elsewhere.

Use the assessment to choose the next result

Hypothetical assessment plan: an operations team wants to evaluate an agent that maintains an internal escalation roster. Staff can invoke it through an interactive client; a scheduled job uses a separate execution identity and configuration. The sponsor wants a single Level 4 claim for both paths. No roster is changed and no test is run for this example.

The following are conditional assessment rules for that proposed review, not collected evidence or assigned levels:

Use the assessment to choose the next result
If the review can establish this conditionDefensible conclusionNext result to establish
Only the interactive path's material usage, configuration, and inventory boundary are establishedLevel 1 may be supported for that boundary; later requirements remain openRelatable observations of the activity needed for the Level 2 review
Policy evidence defines permitted roster changes, but runtime behavior has not been verifiedThe policy supports part of Level 3; it does not establish Level 4 or the other supporting requirementsEvidence that the applicable decision affects the named roster operation on the controlled path
The scheduled identity, configuration, and affected records have not been examinedThat path is unassessed; the interactive finding does not transferA defined scope, identity, authority, and owner for assessing the scheduled path

This produces different work items rather than an average level. The interactive path needs evidence for its next unsupported requirement. The scheduled path first needs an assessment boundary. Buying the same control for both would not resolve the missing knowledge about the scheduled identity or establish that either installation applies it correctly.

Choose the order by the consequences of the operation and the uncertainty that prevents a responsible decision. If the scheduled job can alter records that must remain within a team's authority, resolving that scope may be more urgent than improving a report for the interactive path. This is a conditional priority judgment, not an observed risk finding or a numerical score.

Keep the level subordinate to the operating decision

A higher level does not authorize a new task, resource, or execution identity. A lower level does not, by itself, determine whether a bounded use is acceptable. The accountable owners must establish the controls required for that use and decide which remaining risks or restrictions they can accept. A required preventive control cannot be replaced by a favorable maturity label.

Use the coding agent rollout guide when the next improvement needs a pilot with explicit acceptance criteria. Keep the assessment focused on the outcome that pilot must demonstrate, instead of reproducing its implementation checklist as a scorecard.

Reopen the relevant finding when evidence expires, a configuration changes, a new path is introduced, or an operating gap undermines the earlier conclusion. Preserve findings that remain supported within their original scope. The useful output is a current account of what is demonstrated, what is uncertain, and who owns the next verifiable improvement.