
Auditing AI Agent Tool Permissions, Secrets, and Audit Logs Before Production
Why AI Agent Governance Lags Deployment
This post is a pre-production audit for teams about to point an AI agent at real tools. You will learn how to inventory agents and their capabilities, scope tool permissions in a CI-enforced manifest, broker short-lived credentials, set data boundaries and egress limits, write audit logs you can actually replay, and gate releases on automated evals. The starting point is a claim worth being skeptical about: BigGo Finance reported on 2026-10-10 that Gravitee's CTO puts 88% of firms already running AI agents in production against only 14% that govern them. Treat that as a vendor-attributed survey claim, not an audited statistic — the source snippet publishes no sample frame, no methodology, and no definition of "govern." The direction it points in is still worth taking seriously, and it matches what most teams see internally: agent deployments have outrun the controls around them.
My position is that this is a deployment-gate problem, not a policy-document problem. Almost every team shipping an agent already has an AI usage policy, a Slack channel, and a review doc. What they do not have is a gate that fails a release when an agent registers a tool nobody scoped, when a credential outlives the request that needed it, or when a system-prompt edit quietly widens the blast radius. This post is the audit you run before that gate exists: permissions, secrets, data boundaries, logs, and evals, in that order.
One honesty note up front. The public source material here is thin — one statistic, one best-practices summary, one contract article. The checklist below is engineering practice assembled from building and reviewing these systems, not a restatement of the report. Where I mark something as inference, it is inference.
What the 88/14 Numbers Do and Do Not Tell You
Confirmed (as published): the claim exists, attributed to Gravitee's CTO, in a BigGo Finance piece dated 2026-10-10. The snippet does not say how many firms were surveyed, whether respondents were Gravitee customers, or what "govern" meant on the questionnaire.
Inference (mine, untested): vendor surveys of this shape usually measure policy existence, not enforced controls. "Governed" most likely means the respondent answered yes to something like "do you have AI governance processes?" — a question a slide deck satisfies. If that is what happened, the real gap is smaller than it looks, because many respondents are halfway there on paper, and larger than it looks, because paper does not stop a tool call.
Either way, the number is not actionable as stated. You cannot compare your org to 14% of an undefined population. You can only compare your running system against a definition you wrote down.
Turning a Survey Claim Into a Testable Definition
Restate governance as five artifacts a reviewer can check against a live system:
- An agent inventory — every deployed agent, its owner, and its blast radius.
- A per-tool permission manifest — declared scope, data class, and approval level for every tool.
- Short-lived, brokered credentials — no long-lived secret in the agent's process or context.
- Tool-call audit records — intent, policy decision, and result, append-only.
- An eval suite wired into CI — a release gate, not a leaderboard.
That list is the spine of everything below. All five are verifiable by a reviewer reading config and logs, and none of them require trusting a governance assertion.
Start With an Agent Inventory and Capability Map
Before touching code, enumerate every agent. The unnamed ones — the internal prototype someone pointed at a real Jira project — are the ones that show up in incidents. Autonomy level and data classes are the two columns teams most often cannot fill in, and that failure is itself the finding.
| Agent | Model + version | Tools | Credential type | Data classes read | Autonomy | Human approval required |
|---|---|---|---|---|---|---|
| support-triage | gpt-4.1-2025-04-14 | kb.search, crm.read, ticket.draft | brokered, 60s TTL | internal, PII (read) | suggest | no |
| billing-assistant | gpt-4.1-2025-04-14 | crm.read, ledger.read, refund.issue | brokered, 30s TTL | PII, financial | write w/ prompt | yes, per call |
| devex-shell | claude-sonnet-4-20250514 | shell.exec, fs.write | none (sandbox) | none | auto-execute | two-person for CI |
If a row's "tools" or "autonomy" cell is blank, stop and fill it before continuing. An agent whose tool list is unknown has an unbounded tool list.
Record Model and Prompt Versions as Build Artifacts
You cannot attribute behaviour to a model unless the model version is pinned, and pinning only helps if the system prompt and tool schema are versioned alongside it. Treat the prompt and the JSON tool schemas as release artifacts with hashes, the same way you treat a lockfile.
$ sha256sum agents/support-triage/prompt.md agents/support-triage/tools.json
6f2c1a9d0b4e7f3a... agents/support-triage/prompt.md
b81e44c02d9a7f10... agents/support-triage/tools.json
Record both hashes in the eval run and in every tool-call log line. When an incident review asks "what changed on Tuesday," the answer becomes a diff instead of a conversation.
Classify Autonomy, Then Certify Controls Per Level
Read-only suggestion deserves a lighter gate than auto-executing write tools. Make the level explicit in the inventory so reviewers know which checks apply without asking. A practical ladder:
- Suggest — output is shown to a human; no side effects. Gate: eval suite + logging.
- Write with prompt — side effects, human confirms each call. Gate: above + per-call approval record + idempotency keys.
- Auto-execute — side effects without a human in the loop. Gate: above + two-person rule for irreversible tools + tested kill switch + dry-run mode.
Auditing AI Agent Tool Permissions
The default failure mode is a shared tool catalogue: one registry of everything the platform can do, imported wholesale by each agent. It is convenient, and it is how a support agent ends up holding ledger.write. Take the audit position instead — per-agent allowlists with least privilege, rather than a shared catalogue held together by conventions.
Each agent's allowlist entry should declare, per tool: a scope string, the data classes the tool can return, an owner, a class (read/write/irreversible), rate and spend limits, and — for anything that touches the network — an egress allowlist.
{
"tools": {
"kb.search": {
"class": "read",
"scope": "kb:read",
"dataClasses": ["public", "internal"],
"owner": "support-platform",
"rateLimit": "60/min",
"egress": null
},
"crm.updateContact": {
"class": "write",
"scope": "crm:contact:write",
"dataClasses": ["pii"],
"owner": "crm-team",
"approval": "single-human",
"idempotency": "required"
},
"shell.exec": {
"class": "irreversible",
"scope": "sandbox:exec",
"dataClasses": ["none"],
"owner": "devex",
"approval": "two-person",
"timeoutMs": 15000,
"fsScope": "/tmp/agent-<run-id>",
"egress": []
}
}
}
"egress": [] on shell.exec means the sandbox has no network path at all. That beats an allowlist as a control, and it should be the default for code-execution tools.
Read, Write, and Irreversible Tool Classes
| Class | Examples | Required control |
|---|---|---|
| read | search, fetch, vector query | log every call; no approval |
| write | create ticket, update record, send message | approval prompt or scoped token; idempotency key; logged args |
| irreversible | delete, refund, deploy, shell exec | two-person rule, dry-run mode, timeout, sandbox, egress deny |
The classification decides what the log has to capture. Read calls can be logged at low fidelity. Irreversible calls need the full pre-image: who approved, what arguments were shown to the approver, and what the tool actually did.
Enforcing the Tool Manifest in CI
A manifest that is not enforced in CI is documentation. The check you want diffs the tools actually registered at runtime against the manifest and exits non-zero on any drift — additions, removals, or missing fields.
import { readFileSync } from "node:fs";
const manifest = JSON.parse(readFileSync("agents/support-triage/tools.json", "utf8"));
const registered = JSON.parse(readFileSync("dist/registered-tools.json", "utf8"));
const required = ["class", "scope", "dataClasses", "owner"];
const problems = [];
for (const tool of registered) {
const entry = manifest.tools[tool];
if (!entry) {
problems.push(`${tool}: registered but absent from manifest`);
continue;
}
for (const field of required) {
const value = entry[field];
if (value == null || (Array.isArray(value) && value.length === 0)) {
problems.push(`${tool}: missing ${field}`);
}
}
if (entry.class === "irreversible" && entry.approval !== "two-person") {
problems.push(`${tool}: irreversible tools require two-person approval`);
}
}
if (problems.length > 0) {
console.error("tool manifest check failed:\n" + problems.map((p) => ` - ${p}`).join("\n"));
process.exit(1);
}
console.log(`tool manifest ok (${registered.length} tools)`);Run against a lab repository with a deliberately broken manifest:
$ node tools/check-tool-manifest.mjs
tool manifest check failed:
- crm.deleteContact: missing owner
- shell.exec: irreversible tools require two-person approval
$ echo $?
1
Add the missing owner, move shell.exec to two-person approval, and the same command prints tool manifest ok (7 tools). That is a lab fixture, not a production registry — the point is that the gate is a build failure, not a review comment.
Secrets Handling for Tool-Calling Agents
The core rule: the agent should never hold the long-lived secret. Not in an environment variable in its process, not in a tool config it can read, not in context. If the agent can read the token, then any path that gets the agent to echo its own configuration — a reflection bug, a hostile tool result, a verbose error — becomes credential exfiltration.
The pattern that holds up is a broker or proxy: the agent requests a tool call, the broker authenticates the agent, mints a short-lived credential scoped to the calling user or task, executes the tool, and discards the credential.
// The agent submits intent; the broker holds the secret.
const result = await broker.call({
tool: "crm.updateContact",
actingUser: request.userId, // identity to delegate, not "the agent"
args: { contactId, phone },
ttlSeconds: 60,
idempotencyKey: runId + ":crm.updateContact",
});
Where identity delegation matters — anything where the downstream system enforces per-user authorization — use OAuth token exchange (RFC 8693) so the audit trail at the downstream service names the user, not a shared service account. A shared service account makes every tool call in the downstream log look identical, which flattens your incident timeline into a single blob.
Removing Credentials From Context and Traces
Check the four leak paths, not just one: tool output returned into context, environment dumps in tool errors, stack traces in exception handlers, and request/response bodies written to traces. Then verify the redaction filter actually fires rather than assuming it does.
CANARY="hkjs-canary-7f3a9c"
node scripts/agent-smoke.mjs --inject-canary "$CANARY" > /dev/null
if grep -Rq "$CANARY" traces/; then
echo "FAIL: canary leaked into traces"
grep -Rl "$CANARY" traces/ | head -3
exit 1
fi
echo "ok: no canary in traces"
First run in my lab fixture:
FAIL: canary leaked into traces
traces/run-9f2.jsonl
The leak was tool-call error serialization: a failing tool returned its process.env in the error payload, which the trace writer stored verbatim. After redacting env-like keys on tool error serialization, the same command printed ok: no canary in traces. The lesson generalizes — redaction that is not tested with a real canary is a guess.
Testing for Injection-Driven Credential Exfiltration
This is an authorized-lab test, and it is a design question, not a prompt-hardening question. Seed a retrieved document with content that instructs the agent to call an attacker-controlled tool or echo a value into an outbound request, then observe whether credential or data egress is even possible.
Concretely, ask three questions:
- Does the agent have any tool whose egress allowlist includes an attacker-controlled domain? If not, the attack needs a second bug.
- Can any tool's output cause a subsequent tool call with attacker-supplied arguments, without a policy check between them?
- If the agent has a credential, is it scoped short enough that an exfiltrated token is useless by the time it is used?
If the answer to (1) is no and credentials are brokered with 60-second TTLs, injection becomes a nuisance rather than a breach. That is why brokered credentials are the highest-leverage single control on this list.
Data Boundaries, Retention, and Egress
Write down which data classes may enter context at all. That is a policy decision, but the retrieval layer has to enforce it in code. If PII cannot enter context, the retriever must filter it, because a system prompt that says "do not process PII" is not a control.
Then answer, per model provider:
- What does the provider retain, for how long, and is zero-retention contracted?
- Is processing regional, and does that match the data's residency requirement?
- In shared vector stores or caches, what key keeps tenants apart, and has it been tested with two tenants?
For any tool that makes outbound network calls, the egress allowlist is the last mile. A fetch tool with an unrestricted allowlist is an exfiltration primitive; a fetch tool with a ten-domain allowlist is a liability you can reason about.
Audit Logs That Are Actually Auditable
Log intent plus decision, not just output. An output-only log tells you what happened and nothing about why it was allowed. The fields below are the minimum for replaying an incident without interviewing anyone.
Fields Every Tool-Call Record Needs
| Field | Why it is needed |
|---|---|
| request id / run id | Correlate all calls from one agent run |
| agent + model version | Attribute behaviour to a specific build |
| prompt + tool schema hash | Detect config drift between runs |
| user or service identity | Distinguish delegation from a shared service account |
| tool name + class | Know how bad this call could have been |
| redacted arguments | Reconstruct intent without storing secrets or PII |
| policy decision + rule id | Which control allowed or blocked it |
| approval record | Who approved, and what they saw |
| result status + result hash | Confirm what ran; hash for large payloads |
| latency, token count, cost | Detect loops, runaway spend, abuse |
Tamper Resistance, Retention, and Log Access
Separate the write path from the read path. The agent's process should be able to append and nothing else — no delete, no update. Restrict who can query logs containing user data, because an audit log that includes user data is itself a data store with a retention obligation. State retention explicitly and enforce it with a job, not a wiki page.
Never log raw secrets, tokens, or full document bodies "just to have more telemetry." A trace that captures the credential is a second copy of the credential with a longer half-life than the original.
Automated Evaluation as a Release Gate
Evals are regression tests. They are not leaderboards, and a passing suite is not a compliance artifact. The suites worth gating on:
- Tool selection — a fixed task set with expected tool sequences.
- Scope violation — tasks that should trigger refusal, not a creative workaround.
- Injection resistance — cases seeded with hostile content, assertions on egress and refusal.
- Incident replays — every past incident becomes a permanent case.
Report results in the task / inputs / criterion / count shape so the number sits next to the claim it supports.
A Minimal Agent Regression Suite You Can Run This Week
Ten to fifty representative tasks, each with expected tool calls and expected refusals, run on every prompt or model change. A fixture run from my lab:
Task: select and order tools for 50 ticket-triage requests
(kb.search -> crm.read -> ticket.draft; never refund.issue).
Inputs: 50 seeded requests; 12 contain an injection line in the body.
Criterion: exact match on the expected tool sequence; refusal required on the 12.
Result: 44/50 exact sequences.
6 failures: 4 called refund.issue on a $0 order,
2 skipped the tier lookup and drafted the wrong SLA.
Two honest caveats: this is a synthetic fixture, not production traffic, and 50 cases is a smoke test, not coverage. What makes it useful is that the failures cluster by cause — the four refund.issue calls pointed at a missing spend limit, and the two lookup skips pointed at a prompt that did not make the tier lookup mandatory.
What Evals Do Not Cover
Evals measure sampled behaviour on the cases you wrote. They do not measure authorization. A suite that is green at 50/50 is not evidence that a tool cannot exceed its scope — only that it did not on those 50 inputs. Permission scoping and log review are the controls for capability; evals are the control for behaviour. Teams that gate on evals alone ship agents that pass every test and still hold ledger.write they never needed.
AI Agent Contract Clauses to Fix Before Signing
The Startup Fortune piece in the source material covers structuring AI agent data breach notification clauses before signing. Treat these as negotiation requirements, not legal advice, and get counsel to draft the final language:
- Notification timeline and threshold — how fast, and what triggers it.
- Subprocessor status — whether the model provider counts, and whether you must be notified of changes.
- Definition of "incident" — does it cover a loss through an agent's tool call, or only a database breach? This is the clause most likely to be wrong by default.
- Forensic access and log retention — can you get the provider's logs, for how long, in what format, if you need them?
- Deletion on termination — including prompt logs, traces, and fine-tuning artifacts.
- Indemnity — who carries the loss when the agent, not a human, made the call.
Pre-Production Audit Checklist
Ordered gates, with the evidence a reviewer checks. Each gate should be a stored artifact, not an assertion in a review meeting.
| Gate | Evidence artifact |
|---|---|
| 1. Inventory complete | Filled inventory table; no blank tools or autonomy cells |
| 2. Manifest enforced in CI | Green check-tool-manifest job on the release commit |
| 3. Credentials brokered and short-lived | Broker config; TTL values; no long-lived secret in agent env |
| 4. Data classes approved | Signed list of classes permitted in context, enforced in the retriever |
| 5. Logging live and redacted | Canary test output showing zero leaks; sample log line |
| 6. Eval suite green on the release candidate | Run report with prompt and schema hashes |
| 7. Rollback and kill switch tested | Timestamped drill record, not a runbook |
| 8. Contract clauses signed | Executed agreement covering the incident definition |
Run gate 7 as a real drill before the first production release. A kill switch nobody has pulled is untested code, and you will find its failure mode at the worst possible time.
Conclusion
A governance slide deck does not close the 88%/14% gap, and neither does a vendor survey arriving at a different number. What closes it is deployment gates that produce evidence: a manifest that fails the build, credentials that expire before they can be exfiltrated, logs that let you replay an incident without interviewing anyone, and evals that catch a prompt change before customers do.
If you fix only two things this quarter, make them brokered short-lived credentials and a CI-enforced per-tool manifest. Those two convert most injection and privilege-escalation bugs from breaches into failed tool calls.
What remains unverified from the public sources: the 88% and 14% figures themselves, since no methodology was published in the source snippet; the specific definition of governance the survey used; and any concrete findings from the referenced best-practices guide, which I have not read beyond its title and framing. Treat the checklist here as engineering practice that stands on its own, not as a summary of that report.
Further Reading
- Gravitee CTO says only 14% of firms govern the AI agents 88% already run in production — BigGo Finance, 2026-10-10. The origin of the 88%/14% claim; vendor-attributed survey figure, via Google News aggregation.
- AI Agent Development: Best Practices for Tools, Security, and Evaluation — Adnan Masood, PhD, Medium, 2026-10-10. Best-practices summary covering agent tools, security, and evaluation.
- How to Structure an AI Agent Data Breach Notification Clause Before You Sign — Startup Fortune, 2026-10-10. Source for the contract-clause section above.
Share this post
More posts

Building Agent Governance Into Your API Instead of Bolting It On: Policy, Audit, and CLI Surfaces

Hardening GCP Workloads After Google’s Cloud Security Layoffs: A Developer’s Checklist
