AI Agent Tool-Call Guardrail with Jev
All templates

AI Agent Tool-Call Guardrail with Jev

Check every tool call your AI agents want to run before it runs. Jev, TypeSafe AI's decision model, judges each call; Xano allows it, blocks it, or holds it for a human reviewer, and logs every decision.

Start building in Xano: Access the prompt at https://go.xano.co/start-xano-with-template-skill and run it here with agent-guardrail-jev as your template.
ClaudeOpenAI CodexCursorVS Code

Copy and paste the prompt in your local coding agent of choice.

Template Details

5 Tables 21 APIs 5 Functions

Integrations

Jev
Template Preview
Client appClient appWeb · mobile · agentagent_guard APIagent_guard APIguard/tool-callguard/tool-call/{tool_call_id}auth APIauth APIloginmereviewerssignupreview APIreview APIagents/{agent_id}agentsagentsdashboardpolicypolicytool-calls+2 moreseed APIseed APIseedBusiness logicBusiness logicAuthenticate agentEvaluate tool callReview tool callRouteDecideDataDataguard_agentguard_policyjev_decisiontool_calluserExternal servicesExternal services Jev

Overview

An approval layer for AI agents: before an agent runs a tool, it asks this Xano backend. Jev, TypeSafe AI's decision model, judges the tool call, and Xano allows it, blocks it, or holds it for a human reviewer, then logs every decision.

What it provides

AI agents that can call tools — look up orders, issue refunds, export data, delete records — will sometimes try to do the wrong thing. A prompt injection talks them into exporting customer emails; an ambiguous request turns "pause my subscription" into subscriptions.cancel. Letting every call through is risky, and asking a human to approve every call defeats the point of the agent. Building the in-between layer yourself means a classifier, thresholds, a review queue, an audit log, and a way for the agent to learn what a human decided.

This template is that layer. The agent sends each proposed tool call to POST /guard/tool-call before running it. Xano asks Jev three typed questions — should this run without a human (safe / needs review / dangerous), is it destructive, and is it what the user actually asked for — and gets back probabilities, not prose. A routing rule you control turns those probabilities into a decision: confident-safe calls run, confident-dangerous calls are blocked, and everything uncertain, destructive, or off-request waits in a review queue. If Jev can't be reached or isn't configured, the guardrail fails closed — the call waits for a person. Reviewers approve or reject in a single-file console, and the agent polls for the outcome.

Repo layout

backend/
  api/
    agent_guard/
      api_group.xs
      guard_tool-call_POST.xs
      guard_tool-call_tool_call_id_GET.xs
    auth/
      api_group.xs
      login_POST.xs
      me_GET.xs
      reviewers_POST.xs
      signup_POST.xs
    review/
      agents_agent_id_PATCH.xs
      agents_GET.xs
      agents_POST.xs
      api_group.xs
      dashboard_GET.xs
      policy_GET.xs
      policy_PATCH.xs
      tool-calls_GET.xs
      tool-calls_simulate_POST.xs
      tool-calls_tool_call_id_GET.xs
      tool-calls_tool_call_id_review_POST.xs
    seed/
      api_group.xs
      seed_POST.xs
  function/
    guard/
      authenticate_agent.xs
      evaluate_tool_call.xs
      review_tool_call.xs
      route.xs
    jev/
      decide.xs
  table/
    guard_agent.xs
    guard_policy.xs
    jev_decision.xs
    tool_call.xs
    user.xs
  workflow_test/
    agent_auth.xs
    fails_closed.xs
    review_flow.xs
    seed_loads_demo.xs
  workspace/
    agent_guardrail.xs

How it works

Every proposed tool call runs one pipeline (function guard/evaluate_tool_call):

  1. Authenticate the agent. The agent sends its key in the X-Agent-Key header. Unknown or disabled agents are refused before anything is recorded (guard/authenticate_agent).
  2. Record the proposal. A tool_call row is written as pending_review, so nothing is ever "in flight" without a record.
  3. Ask Jev (jev/decide). One POST https://api.typesafe.ai/v1/systemone call with the agent, the user's request, and the proposed call as state, and three questions: a Choice verdict (safe / needs_review / dangerous) and two Nouls, destructive and matches_request. The call never throws: a missing key, a network failure, or a documented error status (401, 422, 429, 529) comes back as ok: false with a plain-language reason.
  4. Route (guard/route, pure logic). With the active guard_policy thresholds:
    • auto_block — Jev says dangerous with confidence ≥ block_min_confidence (default 0.9).
    • auto_allow — Jev says safe with confidence ≥ allow_min_confidence (0.9), and destructive < max_destructive (0.5), and matches_request ≥ min_matches_request (0.5).
    • human_review — everything else, including any call where Jev did not answer (fail closed).
  5. Log and settle. Jev's answers, latency, token count, HTTP status, and the route go to jev_decision; the tool call becomes allowed, blocked, or stays pending_review.
  6. Human review. A reviewer approves or rejects a held call (guard/review_tool_call), with a note. The agent polls GET /guard/tool-call/{id} and gets may_run: true only for allowed or approved calls.

Thresholds live in the guard_policy table, so tightening or loosening the guardrail is a PATCH /policy (or the console's Policy tab), not a redeploy. Every Jev answer is kept in jev_decision, which over time becomes a labeled record of what your agents tried to do and what people decided.

Common use cases

  • Customer-support agents with write access. A support agent can look up orders on its own, but refunds, account changes, and cancellations are held for a support lead — and prompt-injected requests to export customer data are blocked.
  • Internal ops agents. An agent that retries jobs or cleans up data runs read-only and low-impact work immediately, while bulk deletes and anything touching production data wait for an engineer.
  • Billing and finance agents. Money-moving tool calls (refunds, credits, payouts) get a human check by default; you lower the bar for specific low-risk cases by tuning the policy thresholds as you gain confidence.
  • A shared guardrail for several agents. Each agent registers with its own key, so the review queue and the decision log show which agent proposed what, and you can disable one agent without touching the others.

Quick start

Start with your coding agent. Paste the prompt below into a local MCP-capable coding agent — Claude Code, Cursor, Windsurf, and others. It installs the Xano CLI, connects Xano, and imports this template.

Start building in Xano: Access the prompt at https://go.xano.co/start-xano-with-template-skill and run it here with agent-guardrail-jev as your template.

The agent will walk you through the setup below.

  1. Push the backend to a Xano workspace: xano workspace push -w <workspace_id> -d backend. Xano assigns each API group its own api:<canonical>; find them in the workspace's API groups (or in canonical = "…" after xano workspace pull).
  2. Load the demo data with one call: curl -X POST "https://<your-instance>.xano.io/api:<seed-canonical>/seed". This creates the default policy, three agents, eight sample tool calls, and a demo reviewer (riley.morgan@guard.demo / DemoPass1).
  3. Open the review console — frontend/index.html (replace the four __CANON_<GROUP>__ placeholders with your canonicals if your installer didn't), enter your instance URL, and sign in as the demo reviewer.
  4. Connect Jev. Get an API key from TypeSafe AI (docs.typesafe.ai) and set it as the workspace environment variable JEV_API_KEY. Until you do, the guardrail runs but fails closed: every tool call goes to the review queue with the reason "Jev isn't configured yet".
  5. Wire your agent. Register it in the console's Agents tab (or POST /agents), store the key it returns, and have the agent call POST /guard/tool-call with X-Agent-Key: <key> before every tool call. Run the tool only when status is allowed; when it's pending_review, poll GET /guard/tool-call/{tool_call_id} until may_run is true or the call is rejected.
  6. Before going live, change or delete the demo reviewer, and remove or protect the Seed API group.

API surface

AgentGuard — agent-facing; every call sends the agent's key in the X-Agent-Key header.

Method Path What it does
POST /guard/tool-call Propose a tool call (user_request, tool, arguments). Returns status (allowed / blocked / pending_review), route, reason, and Jev's answers.
GET /guard/tool-call/{tool_call_id} Current status of a call this agent proposed, with the reviewer's note and may_run.

Review — reviewers, Authorization: Bearer <token> from /login.

Method Path What it does
GET /dashboard Counts by status, Jev evaluations and fail-closed count, average Jev latency, and the latest calls.
GET /tool-calls Calls newest first; status=pending_review is the review queue.
GET /tool-calls/{tool_call_id} One call with its logged Jev decision.
POST /tool-calls/{tool_call_id}/review decision approve or reject, with an optional note.
POST /tool-calls/simulate Send a test call as a registered agent through the real pipeline.
GET /policy The active thresholds.
PATCH /policy Change any threshold (each 0 to 1); fields you don't send are unchanged.
GET /agents Registered agents (keys are never listed).
POST /agents Register an agent; its key is returned once.
PATCH /agents/{agent_id} Enable or disable an agent, or rotate its key (new key returned once).

Auth

Method Path What it does
POST /login Reviewer login; returns a 24-hour token.
POST /signup Creates the first reviewer only; closed once any reviewer exists.
GET /me The signed-in reviewer.
POST /reviewers A signed-in reviewer invites another reviewer.

Seed

Method Path What it does
POST /seed Load the demo data (idempotent).

Database Tables

  • tool_call — every tool call an agent proposed: the user's request, the tool and arguments, the status, the route and its reason, and who reviewed it.
  • jev_decision — one row per Jev evaluation: verdict, confidence, probabilities, destructive and matches-request probabilities, latency, input tokens, HTTP status, and the route Xano chose.
  • guard_policy — the confidence thresholds the routing rule applies; only the active row is used.
  • guard_agent — registered agents, each with its own X-Agent-Key and an active flag.
  • user — reviewers who sign in to the console.

Testing

The tests ship in backend/ and run with xano unit_test run_all and xano workflow_test run_all:

  • Unit tests in guard/route (9) cover every routing branch, including stricter thresholds, unknown probabilities, and failing closed. Unit tests in jev/decide (5) mock the Jev call with the response shapes from the TypeSafe API reference — a Choice answer plus two Noul answers, and the documented 401 and 529 errors — and check the no-key path.
  • Workflow tests (4): the seed loads and is idempotent; a held call is approved and rejected by a reviewer and can't be reviewed twice; agent keys are enforced (unknown, empty, and disabled keys are refused); and a full pipeline run records the call and its decision and, with no JEV_API_KEY, holds the call for review.

None of these need a Jev key — the Jev call is mocked from the documented contract, so they prove the integration is correct against the docs, not against the live service. Two live tests in live-tests/ call the real Jev API (an order lookup should be allowed; a prompt-injected customer-data export must not be). They aren't pushed with backend/; after setting JEV_API_KEY, copy them into backend/workflow_test/, push, and run xano workflow_test run_all. Jev's answers are probabilistic, so a live result can vary with the model version.

Environment variables

TypeSafe AI (Jev)

  • JEV_API_KEY — your TypeSafe API key, sent as Authorization: Bearer <key> to https://api.typesafe.ai/v1/systemone. Get one from TypeSafe AI. Without it, every tool call fails closed to human review.

Frequently asked questions

How does the guardrail decide whether an AI agent's tool call runs?
Xano sends the user's request and the proposed tool call to Jev with three questions: a Choice verdict (safe, needs_review, dangerous) and two yes/no probabilities, destructive and matches_request. The guard/route function applies the active guard_policy thresholds: a confident dangerous verdict is blocked, a confident safe verdict that is not destructive and matches the request is allowed, and everything else waits for a human reviewer.
The guardrail fails closed. jev/decide never throws; a missing key, a network error, or a documented error status (401, 422, 429, 529) returns ok: false with a plain-language reason, and guard/route sends the call to human review. Nothing runs unreviewed because Jev was unavailable.
POST /guard/tool-call returns a tool_call_id. While the status is pending_review, the agent polls GET /guard/tool-call/{tool_call_id} with its X-Agent-Key; the response carries the status, the reviewer's note, and may_run, which is true only for allowed or approved calls.
Yes. The thresholds — minimum confidence to allow, minimum confidence to block, maximum destructive probability, and minimum matches-request probability — live in the guard_policy table. Change them with PATCH /policy or the console's Policy tab, and the next tool call uses the new values.
Each agent is a row in guard_agent with its own random key, sent in the X-Agent-Key header. Reviewers register agents, disable them, or rotate keys from the console or the Review API, and every tool call and decision records which agent proposed it.
Unit tests mock the Jev call with response shapes from the TypeSafe API reference (a Choice answer, two Noul answers, and the documented 401 and 529 errors), so the integration is verified against the documented contract. The routing rule and the review flow are tested without credentials; two live tests that call the real Jev API ship in live-tests/ for you to run once JEV_API_KEY is set.
Every template runs on any Xano plan, including the free tier. All you need is a Xano account and a coding agent connected to the CLI and Developer MCP.
You can import and run it with your local coding agent — it works against the seed data it ships with. Paste this prompt into your agent: Start building in Xano: Access the prompt at https://go.xano.co/start-xano-with-template-skill and run it here with agent-guardrail-jev as your template. You can also do a manual clone from Github, still using the Xano CLI and Developer MCP to connect to Xano to build and deploy.
Yes — that's the intent. The backend is plain XanoScript that you can read and AI can build on: add or edit tables, endpoints, and logic, or swap integrations. It ships with tests you can extend as you go.
Xano gives you an enterprise-grade backend — database, APIs, and logic — that AI can build and you can visually inspect, edit, and trust. Hosted, scalable, and production-ready from day one. Speed of generation, without the black box.

Get started for free today

Xano gives you everything you need to ship modern applications—fast, securely, and at scale.

Get started