Loading…
The most comprehensive AI red‑teaming platform

We find out what your AI can be made to do.

Chat, RAG, APIs, tools, MCP servers, autonomous agents, voice — if it is part of your AI system, we attack it. Every surface, one engagement, and every finding proved with evidence you can re-run yourself. See what we actually found — five anonymized engagements

Subscription-aware access is loaded automatically.

Responsible disclosure
Zorio Flipkart Swiggy Orbitshift Bolna Fynd Twincafe Apollo 247 Yellow.ai
01 / Coverage

Every surface your AI exposes. All of them, in one engagement.

AI stopped being one chatbot on a website. It is the assistant your customers type to, the knowledge base behind it, the APIs underneath, the tools and MCP servers it calls, the agent working a queue at 3am with real credentials — and increasingly the voice on the end of a phone line. Each fails differently, so each is tested differently. Under one report.

Conversational AISupport bots, in-product copilots, internal assistants. Talked into leaking another customer’s data — or its own instructions.
  • Multi-turn jailbreaks and system-prompt extraction, scored across the whole conversation — an agent that refuses every direct request still gives its rules away a piece at a time, which per-reply scoring never sees
  • Cross-user and cross-tenant data exposure through conversation alone
  • Filter evasion: encoding, unicode, and payloads split across separate requests
  • Indirect injection through documents and images the assistant reads
  • Fabricated policies, prices and entitlements a customer could act on — nothing breached, the answer is simply false
RAG & knowledge retrievalPoisoned context, forged citations, and documents that quietly carry instructions to your model.
  • Indirect prompt injection through retrieved content, proved with a planted canary that shows the embedded instruction actually executed
  • Retrieval steering and citation forgery
  • Supply-chain poisoning of ingested documents
  • Cross-tenant retrieval — reaching documents indexed for another customer
  • Every plant is marked, logged, and handed back in a cleanup manifest naming the exact values
APIs & authorizationWhere AI breaches actually happen. We name the exact victim record we reached, not a theoretical “could happen”.
  • BOLA / IDOR with a cross-tenant differential — the record one identity reached that belongs to another, named in the finding
  • BFLA — role-forbidden functions invoked while holding an ordinary user’s session
  • The owned-record baseline is discovered from the token itself, so nobody enumerates identifiers by hand
  • Token and JWT forgery: alg:none, weak-key signing, missing expiry checks, replay
  • GraphQL resolvers tested individually — one URL is many authorization decisions, and reading only the URL sees one
  • Race and TOCTOU windows on state-changing actions, with before/after state proof
  • Business rules mapped from the system’s own explanations, then broken
Tools & MCP serversA Model Context Protocol server is a first-class target here, not a degraded one.
  • All three primitives — tools, resources and prompts — plus the JSON-RPC layer beneath them
  • Session binding, catalogue drift, server-initiated callbacks, and the OAuth flow in front
  • Poisoned tool metadata: descriptions are handed to a consuming model as trusted context, so any agent that merely lists your server reads them
  • Unauthorized tool calls and parameter smuggling against a REST tool gateway
  • Both transports: streamable HTTP and HTTP+SSE
  • With two accounts we prove cross-tenant access with real record ids rather than guessed ones
Agentic workflowsNo user in the loop, real credentials, real tools. Approval skipped, privilege escalated, state raced — with proof of the change.
  • Multi-step plan manipulation and tool-chain hijacking
  • Approval and workflow step-skipping, captured with before/after state proof
  • Memory poisoning that survives into a later session
  • Cross-session bleed between users
  • Compromised-identity starting points: assume an employee's credentials are already stolen and test what the session can reach - lateral movement and authorisation boundaries, not the login
  • Structured-output and schema confusion at the parser boundary
  • Concurrency races on state-changing actions — double-spend and check-then-use
VoiceA real phone call, or the widget on your site. We speak, listen, interrupt — and prove the agent actually heard us.
  • Real calls over the phone network, or the realtime session behind your web widget
  • Voice-native evasion, barge-in interruption, and the session control plane in front of the agent
  • Every spoken attack is round-tripped through our own speech recognition first: if the target could not have heard it, it is reported undelivered, never as a pass
  • Call recording and transcript attached to every finding
See how voice testing works ›
One engagement covers all six. Voice is a surface in our engine, not a separate product — and the same is true of MCP, RAG and the rest. That is what makes a guardrail which holds in chat and fails on a phone call visible: it becomes a single finding in a single report, and it is invisible to anyone testing the two separately.
02 / Depth

It works out your system first. Then it goes after it.

A scanner runs a list. This does reconnaissance, forms a view of your business logic, and spends the rest of the run acting on what it learned — which is why the second half of an engagement finds things the first half could not have known to look for.

It finds your surface, not your docsThe legacy route nobody documented is tested too.
  • Endpoints come from browser traffic captures, your production JavaScript bundles, gateway logs, historical archives and specs — not just whatever was written down
  • One real engagement surfaced six API hosts the browser already ships, none of them in the documentation
  • Every endpoint records how we found it, so a coverage claim can be checked rather than trusted
  • Anything discovered outside your authorised boundary is listed as a lead we deliberately did not touch
It learns during the runAn identifier seen in one response becomes a probe against a different endpoint.
  • A path leaked in an error message becomes a target enumeration never had
  • Attack families that keep coming back empty are deprioritised; the budget follows what is working
  • A generative strategist proposes the next attack family, vector or chain from what the run has observed - a queue of cold families is pre-empted by fresh ideas
  • Part of the request budget is held back specifically for work the run could not have planned in advance
  • The report separates what we found before the attack from what we found during it — a scanner has nothing in the second column
It chains findings across layersTwo mediums that combine into a critical are reported as the critical.
  • A token disclosed by the model becomes a credential re-used against your API — read-only, to prove the access it unlocks. Session hijack, not just disclosure
  • A prompt injection that reaches a tool that changes state is one chain, not three unrelated rows
  • Every chain ships as a sequence you can replay step by step. A chain a client cannot re-run is a story, not a finding
  • Chains cross modalities too — something learned on a phone call used against the chat API in the same run
It tells you what it could not testSilence and safety look identical on a page, and only one of them is true.
  • A correct refusal is reported as a control that worked, which is worth paying for too
  • An action behind a verification step is reported as not tested — never as clean. Hand us a test account’s authenticator secret and we generate the code and actually test it
  • An attack class that never reached your system — refused tool, unreachable surface, budget cut short — says so instead of showing a clean row
  • Claims are capped by evidence automatically and the report shows the downgrade. A status code alone can never be a Critical
03 / Proof

Findings from live systems, with the evidence attached

Not categories we could test. Things we have actually got out of running AI products.

Voice · instruction disclosure

A voice agent gave up its operating instructions across four spoken turns

Every individual turn was a polite refusal. Judged turn by turn it was a clean pass; judged across the whole conversation it reconstructed the agent's own rules.

evidence: call recording + transcript
Voice · session control plane

A widget's session endpoint minted billable sessions, unauthenticated

No prompt cleverness needed. The HTTP layer in front of an agent is often the softest part of it, and every call provisioned real capacity on the operator's bill.

evidence: request/response pair, replayable
MCP · poisoned tool metadata

Five tool descriptions on a live MCP server carried instructions aimed at the calling agent

Tool descriptions are handed to a consuming model as trusted context, so any agent that merely lists that server reads them. The same engagement found the server's object authorisation to be sound - and the report says so, because a clean result you can trust is worth paying for too.

evidence: request/response pair, replayable
Discovery · front end

Six API hosts the browser already ships, none of them in the docs

Read straight out of the production JavaScript bundle. You cannot test a surface you never found, so discovery caps every other number on this page.

evidence: source reference, per host

And the part most vendors leave out: against a hardened agent, none of our ten extraction techniques recovered a discrete secret. Instruction reconstruction is what actually works, so that is what we claim — and the report states which attacks failed, not only which landed.

04 / Why your pentest missed it

Your pentest never talked to your AI

A scanner looks for flaws in code. The failures that matter in an AI system are behavioural, and they surface only when something argues with your model the way an adversary would.

The surface is a conversation

Vulnerabilities live in prompts, system instructions, and multi-turn context - invisible to any scanner that doesn't talk to your model like an attacker.

Business logic is the bug

Not a buffer overflow - convincing your AI to expose another user's data, skip an approval, or leak a system prompt. That needs domain-aware attacks.

Behaviour drifts every update

A guardrail that held last quarter can break after a fine-tune or prompt change. Point-in-time tests go stale; continuous red teaming catches it.

Regulators now expect it

The EU AI Act, NIST AI RMF and sector guidance increasingly require documented adversarial testing of high-risk AI.

Built for the two teams that answer for AI risk

Teams building AI products

Security stops being the thing that delays the launch and becomes the thing you sell with. Continuous, evidence-grade proof that your AI holds up — ready before the questionnaire arrives.

  • Catch prompt-, API-, tool-, and agent-layer regressions before release
  • Shipping a voice agent? Have it attacked over a real call before your customers - or their attackers - do
  • Run pre-release or in CI/CD - no LLM-security expertise required
  • Answer customer security reviews with mapped, reproducible evidence

Teams securing AI inside the enterprise

Your staff are already using internal AI that touches sensitive data. You need to know it cannot be talked into leaking across users or over-reaching its tools — and to show an auditor you tested for exactly that.

  • Probe internal copilots for cross-user data exposure and prompt leakage
  • Findings in the evidence format your security programme already uses
  • Documented proof mapped to the EU AI Act, NIST AI RMF, ISO 42001, GDPR & DPDPA
05 / The detail behind it

The three things enterprise buyers ask about

The full attack catalogue, how defensible the output is under audit, and how an autonomous attack engine is kept from being misused. Open whichever one your review board cares about.

Pick one to explore - the detail below updates

Jailbreaking the model is table stakes. Real AI breaches happen where the model meets your APIs, your tools and your business workflow — so every family below is exercised against whichever of your six surfaces actually exposes it, and proved with evidence rather than theory.

Delivered over whichever channel your users actually get

All four layers are attacked over text and over live audio in the same engagement, so a weakness that only appears when the channel is speech still lands in the same report. See how voice testing works ›

Attack families

25 attack families plus an always-on credential-reuse pivot, each mapped to OWASP and MITRE ATLAS. Every one is exercised against whichever surfaces your system actually exposes.

See all 26 families and their mappings
Prompt injection - direct

Direct instruction-override against the system prompt.

OWASP LLM01 · AML.T0051
Prompt injection - indirect

Injected instructions via retrieved content, docs, and images.

OWASP LLM01 · AML.T0051
Data exfiltration

Extracting secrets, system prompts, and PII from the model.

OWASP LLM02 · AML.T0057
Authorization bypass

Role escalation, approval spoofing, and authority abuse.

OWASP LLM03 · ASI03
Tool abuse

Unauthorized tool calls and parameter smuggling.

OWASP LLM03 · ASI02
RAG attacks

Retrieval steering, context poisoning, and citation forgery.

OWASP LLM01 / LLM09 · ASI06
Vector & embedding weakness

Embedding-space geometry attacks - cross-tenant inference, semantic-cache poisoning, retrieval jamming, membership inference.

OWASP LLM09 · ASI06
Obfuscation

Encoding, unicode, and multi-step payload hiding to evade filters.

OWASP LLM01 · Evasion
Multimodal injection

Image- and document-borne instruction attacks.

OWASP LLM01 · Multimodal
Agentic planning bypass

Plan and tool-chain manipulation against agentic systems.

OWASP LLM01 / LLM03 · ASI01
Memory poisoning

Long-horizon state corruption and memory steering.

OWASP LLM01 / LLM05 · ASI06
Cross-session bleed

Context leakage across users and conversations.

OWASP LLM02 · ASI06
Supply-chain poisoning

Poisoned retrieved content and document supply chain.

OWASP LLM01 / LLM04 · ASI04
Schema confusion

Structured-output and parser-boundary attacks.

OWASP LLM10 · ASI02
API authz - BOLA / BFLA

Object- and function-level authorization probes against your API.

OWASP API01 / API05
Token & JWT forgery

alg:none, weak-key signing, missing expiry checks, and token replay against your auth layer.

OWASP API02 · Broken auth
Concurrency & race (TOCTOU)

Double-spend and check-then-use races on tool/workflow actions, with before/after state proof.

OWASP API · Business logic
Goal hijacking

Redirects the agent's terminal objective toward the attacker's goal while it appears on-task.

OWASP LLM01 · ASI01 · AML.T0081
Inter-agent trust escalation

Self-asserted identity, role and permission claims accepted without verification.

OWASP LLM01 · ASI03 / ASI07 · AML.T0091
Lateral movement through the agent

Chains the agent's own tools to hop across connected systems from an assumed identity.

OWASP LLM03 · ASI03 · AML.T0053
Unsafe output handling

Model output reaching a browser, shell, or query as trusted content - the reply is the payload rather than the leak.

OWASP LLM10 · ASI05
Business-rule abuse

Gets the system to explain the pricing, limit, and eligibility rules behind its answers, then breaks them.

OWASP API · Business logic
Fabricated answers

Policies, prices, or entitlements the system invents and a customer could act on. Nothing is breached - the answer is simply false.

OWASP LLM07 · ASI09
Misinformation / misleading output

Unsupported, false or incomplete claims that a human, workflow or tool acts on - false state, critical omission, forged evidence.

OWASP LLM07 · ASI09
Model extraction

Assesses how far the deployed model can be partially replicated or mined by black-box querying.

OWASP LLM06 · LLM02
Credential-reuse escalation

Harvests a leaked token/key/cookie and re-authenticates with it (read-only) to prove the access it unlocks - session hijack, not just disclosure. Always on.

OWASP API02 · Credential access

It doesn't end at the report. Ship a fix, re-run the engagement, and every finding returns a fixed / still-failing / regressed verdict - so you can prove the gap is actually closed.

06 / Voice

Voice agents, attacked over a real call

Elsewhere, “we support voice” usually means a transcript pasted into a text API. We place the call. We speak, listen, interrupt, and verify the agent actually heard us — same attack families, same evidence standard, same report as every other surface.

Telephony

Over a real phone call

Give us the number your customers call. We dial it, speak to your agent, listen to what it says back, and press the attack across the turns of a live call.

  • Your published number, or a SIP trunk into your stack
  • Twilio and Telnyx media streams supported out of the box
  • Navigates IVR menus and keypad (DTMF) prompts to reach the agent
  • Hard caps on calls, turns and duration - your bill stays predictable
Web & embedded

Through the widget on your site

The "talk to our AI" button on your product or landing page is a live, usually unauthenticated entry point. We join the same realtime session your visitors do.

  • WebRTC and WebSocket audio - LiveKit, OpenAI Realtime, ElevenLabs, Deepgram, or a custom protocol
  • Also attacks the control plane in front of the agent: the endpoint that mints a session and the token that gates it
  • Reads your front-end bundle to find the API hosts and identifiers the browser already ships
Cross-modal

One engagement, both channels

Voice is a surface in our engine, not a separate product - which is what lets the same run cover text and audio together.

  • No second configuration, no second engagement, no separate report
  • Findings stay directly comparable across channels
  • A guardrail that holds in chat and fails on a call is itself the finding - and it is invisible to anyone testing the two separately
What changes when the channel is audio — six voice-specific classes
Cumulative instruction extraction

Scored across the whole conversation, because an agent that refuses every direct request still gives its rules away a piece at a time - which per-reply scoring never sees.

OWASP LLM07 · Disclosure
Voice-native evasion

Text tricks like base64 do not survive being spoken. Paced delivery, spelling-out and phonetic framing do - so those are what we use.

OWASP LLM01 · Evasion
Interruption / barge-in

Talking over the agent mid-sentence to cut off its safety preamble and land the request while it is recovering.

Voice-specific · Turn control
Session control plane

Unauthenticated session minting, JWT scope and expiry tampering, and weak-signing-key recovery on the token that gates the call.

OWASP API02 · Broken auth
Identity & authority pressure

Claiming to be a supervisor, engineer or verified customer on a channel where there is often no authentication at all.

OWASP LLM06 · Excessive agency
Human-transfer boundary

We test where the agent hands off to a person - and stop there. A tripwire ends the call rather than occupying one of your staff.

Rules of engagement

Every spoken attack is round-tripped through our own speech recognition before it counts: if the target could not have heard it, we report it as undelivered, never as a pass. The call recording is attached to the finding — our audio on one channel, your agent on the other.

Reachable over SIP/PSTN for real phone calls, and over WebRTC or WebSocket audio for anything embedded — Twilio, Telnyx, LiveKit, OpenAI Realtime, ElevenLabs, Deepgram, or your own protocol.

See the transport matrix
SIP / PSTNReal phone calls
TwilioMedia Streams
TelnyxMedia Streams
LiveKitWebRTC
OpenAIRealtime API
ElevenLabsConversational AI
DeepgramVoice Agent
CustomRaw WebSocket audio
07 / Questions

Answers to what buyers ask first

What is AI red teaming, and why is it different from a regular pentest?

AI red teaming is adversarial testing specifically designed for LLM-based systems. A regular penetration test looks for code-level vulnerabilities - SQLi, XSS, misconfigurations. AI red teaming looks for behavioral vulnerabilities: can an attacker override your system prompt? Can they access another user's data through the chat interface? Can they manipulate the model into bypassing an approval workflow?

These risks don't appear in CVE databases and can't be detected by any scanner. They require an engine that generates context-aware, business-model-aware attack prompts and evaluates the model's responses the way a skilled human adversary would.

How is this different from model-only red teaming tools?

Most "AI red teaming" tools test foundation models in isolation - jailbreaking GPT or testing Claude for harmful content. That's useful, but it's not the risk your business faces.

HiltLock tests your deployed AI application: your system prompt, your RAG pipeline, your tool integrations, your user identity model, your data access patterns. The attack surface is the full stack, not just the model layer. We generate targeted attacks that incorporate your org context, data domains, and known capabilities - the same information an insider threat or determined external attacker would use.

What does the engagement process look like?

It's fully automated and runs through the dashboard. You configure your target endpoint, describe your organisation and data domains, and select the attack families to test. The engine then runs a four-phase process:

  1. Recon - probes your AI to map capabilities and guardrails
  2. Exploit - runs multi-turn attack chains across all selected threat families
  3. Verify - re-runs successful attacks to confirm reproducibility
  4. Judge - each finding scored for severity, with evidence verified against the actual transcript (no hallucinated quotes)
What do I get at the end of a run?

A structured report containing:

  • Full finding report - every confirmed vulnerability with description, severity, and remediation guidance
  • Confirmed evidence - the exact conversation traces that produced each finding, reproducible on demand
  • Standards & regulatory mapping - each finding mapped to the OWASP LLM / API / Agentic Top 10 and MITRE ATLAS, plus article-level exposure under GDPR, the EU AI Act, India's DPDPA, NIST AI RMF, ISO/IEC 42001, and OECD AI Principles - with version-tracked citations
  • Regression re-testing - re-run after a fix and every finding returns a fixed / still-failing / regressed verdict
Is it safe to point this at our production system?

Yes - restraint is built in, not bolted on. The engine runs only against targets you have verified you control: access is default-deny by domain, and a new target must be explicitly authorized before any run can start, so it can never be aimed at someone else's AI.

In production it is non-destructive by design - it never issues PUT / PATCH / DELETE requests or destructive tool/MCP actions, and declarative rules-of-engagement are enforced uniformly across every executor. An always-on SSRF guard refuses internal, loopback, and cloud-metadata addresses, and each run stays confined to the endpoint you provide. Secrets and PII are redacted before anything is stored. Many customers still prefer to start against a staging environment - that works too.

Do I need to install anything or expose my system?
No agents, SDKs, or code changes are required. The engine communicates with your AI system over HTTPS exactly as a real user would. You only need to provide the endpoint URL and an API key or bearer token. Your system does not need to be publicly accessible - you can use a staging environment or allow-list the engine's IP range.

To make setup fast even for non-technical users, the configuration screen can auto-fill itself from a browser capture: paste a cURL command, or upload a .har file, from a single real chat with your AI, and the system URL, connection details, request format, and login token are detected and filled in for you. The file is parsed entirely in your browser and never uploaded.

Can you test our voice agent - and what do you need from us?

Yes, and it is one of the few areas where automated adversarial testing barely exists today. There are two routes in, and we support both:

  • By phone - give us the number, a window during which test calls are acceptable, and (optionally) a SIP trunk if you would rather we not touch the public number. We dial in and speak to the agent.
  • By web or embedded widget - give us the page carrying the voice widget. Exactly as with chat, a .har capture of one real conversation auto-fills the connection details, and is parsed in your browser and never uploaded. Otherwise we need the realtime host, the endpoint that mints a session, and the audio protocol.

Two other things help: a call budget (each attack chain is a real call, with real speech-processing cost on your side - we cap calls, turns and duration, and the defaults are deliberately conservative), and the phrase or number your agent uses to transfer to a human, so our tripwire ends the call rather than occupying one of your staff.

To be explicit about what we do not do: we never clone a real person's voice or impersonate a named individual. Every call is made in our own clearly synthetic voice. The attack is on your agent's reasoning and its authorization boundaries, not on your callers.

On the recordings themselves: they are encrypted, never public, reachable only through an authenticated request scoped to your run, and they expire automatically — 90 days where the call produced a finding, 7 days otherwise. The transcript and the finding are retained, so the evidence trail outlives the audio.

How long does a run take?
Runtime depends on the number of attack families selected, the recon depth, and your AI system's response latency. A focused single-family run typically completes in 1-2 hours. A full-spectrum run across all attack families usually takes 12-24 hours. Results are available in the dashboard as soon as each phase completes.
Can I start in demo mode first?
Yes. You can explore the full dashboard configuration interface in demo mode without a subscription. To execute a live run against your AI system and access reports, subscribe through AWS Marketplace. Subscriptions are usage-based with no long-term commitment.