LLM penetration testing services, compared by scope and price

By The PenTest Index · Offers checked October 10, 2026

LLM penetration testing services attack your AI feature to show what it can be tricked into leaking or doing. Five providers publish a price. For a document assistant with one or two tools, Invadel lists $7,500 and Software Secured starts at $10,800; Schellman's minimum is $16,000. The price follows what your AI can read and do, so match scope before comparing numbers.

One warning before the tables: "AI pentest" is sold with three different meanings, and only one of them tests your LLM feature. We read the service and pricing pages of 15 providers and sorted them below. We did not buy or run any of these tests.

Is this a test of your AI, or a test done by AI?

Only the first row below is what you were asked for. The other two get sold under the same words.

What you are buyingWhat happensExamples we readIs it the right purchase?
A test of your LLM featureTesters attack your chatbot, the data it reads and the actions it can takeInvadel AI/LLM penetration testing, Schellman AI Red Teaming, Blaze AI penetration testingYes
A normal pentest run by AIAn AI agent tests an ordinary web app or networkCobalt Autonomous Pentest, Synack Sara Pentest, Astra Pentest AutoNot for this job, unless the provider confirms in writing that it covers your LLM feature
A tool you run yourselfSoftware fires test prompts at your modelPromptfoo, garak, PyRITUseful between tests. It gives you no independent report

The middle row is the easy one to buy by mistake. Synack's pricing table lists Sara's asset type as "External web or host" (Synack pricing, checked October 10, 2026). Cobalt's Autonomous Pentest page names no LLM scope and says it "does not produce compliance attestation reports" (Cobalt, checked October 10, 2026). Both companies sell a separate LLM test. Astra lists Pentest Auto as autonomous testing for web and SaaS apps (checked October 7, 2026).

If you wanted an AI tool to pentest your regular app, that is a different purchase: see automated penetration testing.

What do LLM penetration testing services actually test?

They test the application you built: the prompt, the data it can reach, the actions it can take and what happens to its answers. Some services also test the model itself when that is in scope. Cybersecify puts it plainly: "We test your application, not the foundation model."

What your AI feature can do decides how big the test is.

Your AI featurePlain meaningWhat a tester should tryWhat to tell providers
Chat onlyIt answers from fixed instructionsMaking it ignore its rules; pulling out its hidden instructions; getting unsafe output onto the pageWhere the chat lives; public or logged-in
Reads your documents (often called RAG)It looks things up in your data before answeringGetting one customer's data from another customer's account; planting instructions inside a document it will readWhat data it reads; how customers are kept apart; who can add documents
Takes actions (an agent with tools)It can send, change, refund or deleteTricking it into an action the user is not allowed to takeEvery tool it can call and what each can change
Connects to outside tools or other agents (MCP servers, plugins)It trusts software you do not runA poisoned tool steering the agentEach connector and who runs it

Think of a helpful librarian. The risk is not that the librarian says something rude. The risk is that a convincing question gets another customer's file handed over. Once the librarian can also issue refunds, there is a till to protect as well.

Two things to settle in writing. First, cost and uptime attacks are sometimes left out: Schellman says it "does not test for token exhaustion attacks unless specifically requested." Second, a security test is not an accuracy review. It will not tell you whether the answers are right.

Most providers map findings to the OWASP Top 10 for LLM Applications. OWASP published a 2026 edition in August 2026, and it reordered and renamed entries. Excessive Agency is now LLM03, and there is a new LLM08, Hidden Context Exposure (OWASP's 2026 list, checked October 10, 2026). Some provider pages still cite the 2025 edition. Ask which edition your report will use. OWASP's list is guidance on risks. It is not a certificate, and we are not affiliated with OWASP.

Which LLM penetration testing services publish a price?

Five of the 15 providers we read publish a price that applies to LLM testing. They are not the same purchase. One is per scope, one is per year, three are per engagement, and two of those are starting figures.

Every cell is provider-published and was read on the provider's own page on October 10, 2026. None is a quote for your project.

Provider and offerPublished priceWhat that price covers, as statedRetest, as statedStated timingConfirm before you rely on it
Astra, Pentest Expert$5,999 per year for one targetTarget types listed as "Web, Mobile App, Cloud, Network, AI, MCP etc."; includes "Pentest of AI components within target scope"2 manual re-scans; request within 30 days of findings being reported, after fixing at least 50% of Critical/High vulnerabilities10 to 20 working days for the manual test, depending on which Astra page you readWhether your LLM feature counts as its own target or as part of your web app target
Cybersecify, Startup planINR 74,999 plus taxes for one scope (about $790; the rupee price is binding). A second Startup scope is INR 44,999 (about $470; maximum two scopes). Its Growth plan is INR 179,999 plus taxes for two scopes with additional audit evidenceOne AI scope: "up to 8 agent tools, up to 7 output sinks, one model endpoint". "An agent and the API it calls are two scopes"One full retest within one month of the first report, no extra charge5 business days for one scope, 10 for twoHow many scopes your feature is; whether your Startup quote includes a customer-facing letter (the published letter is Growth-only)
Invadel, AI/LLM penetration testingSmall $4,500 · Medium $7,500 · Large $12,000+, each labeled "Fixed price". The page shows a $ sign and does not name the currencySmall: one feature, "a single model, basic guardrails, and no tool access". Medium: "a retrieval (RAG) pipeline, several prompts, and one or two tools". Large: many tools, agents, MCP servers"Free retest included"; no count or deadline givenTesting "typically begins within a week of scoping"; about a week for one featureThe retest deadline; which tier your feature lands in
Schellman, AI Red Teaming"No less than $16,000 for a single AI red team engagement""LLMs and RAG implementations"; prompt injection, data leakage, output handlingNot stated"Around 2 weeks"Retest terms; start date; whether your agent's tools are in scope
Software Secured, AI Pentesting"Starts at $10,800 USD"Model behavior, document retrieval, agent workflows, connected tools"3 rounds over 12 months"Scheduling "within 3-6 weeks"; report 48 to 72 hours after testingWhat the starting price includes for your feature; your actual start date

Sources: Astra pricing · Cybersecify pricing and service page · Invadel pricing · Schellman · Software Secured pricing and service page.

Three things this table shows that a single price does not.

The cheapest tier may not be your test. Invadel's $4,500 tier is defined as "no tool access." If your assistant reads documents and calls a tool, you are in the $7,500 tier. That is a $3,000 difference decided by one fact about your product.

A scope is not always what you think. Cybersecify says an agent and the API it calls count as two scopes. For an assistant that calls tools, we work that out on the two-scope Startup plan as INR 74,999 + INR 44,999 = INR 119,998 plus taxes, or about $1,260 on Cybersecify's own indicative dollar figures. Its Growth plan costs INR 179,999 plus taxes for two scopes and adds compliance mapping and other testing. Confirm the scope count and plan before you budget on $790.

Retest deadlines differ by months. Astra and Cybersecify close the retest in about a month. Software Secured gives 12 months. Invadel and Schellman do not state a deadline on the pages we read.

We do not publish a "typical range" for this kind of test. Most providers quote, and a range built from five different units would mislead you.

Which offers deserve a closer look?

Start with the thing that would rule an offer out. For most buyers that is the retest deadline or the scope, not the price.

Your situationLook first atThe condition that matters
A document assistant with one or two tools, and you want a published fixed priceInvadel Medium ($7,500)Ask for the retest deadline in writing
Your fixes will take longer than a monthSoftware Secured (from $10,800; 3 retest rounds over 12 months)Stated scheduling is 3 to 6 weeks, so ask for a start date
A small feature, a small budget, and you can fix findings inside 30 daysAstra Pentest Expert ($5,999 per year) or Cybersecify (from about $790 per scope)Both retest windows close in about a month
An agent that can move money or data, or many connected toolsInvadel Large ($12,000+), Schellman (from $16,000), or a quote from Bishop Fox, Blaze or HackerOneGet every tool named in the written scope
You train or host your own modelTrail of Bits, NetSPI's customized testing, Bishop FoxThis is deeper work than an application test; expect a quote
You already have a pentest providerYour current providerAsk what they would add for the AI feature, and the price

That last row may save you the whole purchase. A web app test usually covers the login, the pages and the API around your chat. It often does not try to trick the model unless you ask. Send your provider the scope lines further down and ask one question: "Does our current scope include prompt injection, document access between customers and the assistant's tool calls? If not, what would it cost to add them?" See what a standard web application penetration test covers.

If you already know your scope, go straight to the provider.

See Invadel's AI/LLM tiers and what each includes

See Software Secured's AI pentesting prices

See what Astra Pentest Expert includes

See Cybersecify's plans and scope rules

See Schellman's AI Red Teaming service

The AI feature is one part of your scope. Providers will also ask about user roles, environments, your deadline and who reads the report. Find My PenTest Match walks you through those and gives you a scope checklist to copy or print. It is free, asks for no email, and sends nothing to providers. Choose "A web app and its API", since that is where most LLM features live, then add the AI lines from this page.

Find My PenTest Match

A worked example: one buyer, five published offers

For this buyer, two offers are worth the first calls: Invadel Medium and Software Secured. Each has one open question. Here is how we got there.

The buyer is made up. Say you run a 30-person SaaS company. Your support assistant reads customer documents and can call two read-only tools: look up an order and look up a ticket. A large customer's security questionnaire asks for a penetration test of the AI feature. They want the report in six weeks. Your team needs about 45 days to fix what turns up, and the customer wants proof that the fixes were rechecked.

That gives four requirements. The test must cover document retrieval and the two tools. It must check that one customer cannot pull another customer's documents. The report must arrive within six weeks. A retest must still be available on day 45 after the findings.

This is the PenTest Index Purchase Check: we take one buyer's requirements and hold each published offer up against them. "Supported" means the published terms support that one requirement. It says nothing about how good the testing is.

OfferRequirementFindingWhyQuestion to send
Invadel Small, $4,500Covers retrieval and toolsMismatchThe tier is defined as "no tool access"None. Use Medium
Invadel Medium, $7,500Covers retrieval and toolsSupportedThe tier names a retrieval pipeline and "one or two tools""Please confirm both tools and cross-customer document tests are in the Medium scope."
Invadel MediumRetest on day 45UnresolvedA free retest is stated; no deadline is given"Until what date can we request the free retest?"
Software Secured, from $10,800Retest on day 45Supported3 retest rounds over 12 monthsNone on this point
Software SecuredReport within six weeksUnresolved, at riskStated scheduling is 3 to 6 weeks before testing starts. Only the early end leaves room for the test and report"Can you start within two weeks? Please give the report date in writing."
Astra Pentest Expert, $5,999 per yearRetest on day 45MismatchRe-scans must be requested within 30 days of findings being reported"Can you include a re-scan requested on day 45, and at what cost?"
Cybersecify, about $1,260 for two Startup scopesRetest on day 45MismatchOne retest within one month of the first report"Can the retest window be extended to 45 days?"
Schellman, from $16,000Retest on day 45UnresolvedNo retest terms on the page we read"Is a retest included, how many, and until when?"
All fiveCross-customer document testUnresolvedNo service page can promise what your contract will say"Will you test document access between two test customers, and show the evidence in the report?"

What this buyer should do. Call Invadel and Software Secured first. If Invadel's retest deadline covers day 45, it is the lower published price for this scope. If Software Secured can start within two weeks, it is the one whose published retest terms already fit. Astra and Cybersecify cost less, but both retest windows close before this buyer's fixes are ready, so they only work with a written extension. Schellman may fit at a higher minimum once its retest terms are known.

A mismatch here rules an offer out for this buyer only. A team that fixes findings in two weeks would read the same table and reach a different answer.

These are our readings of published terms. No provider has quoted for this example.

Who else offers LLM penetration testing?

Ten more providers describe a relevant service and ask you to request a quote. This table shows what each page says and what it leaves open. All cells are provider-published, read October 10, 2026. Providers are listed A to Z; this is not a ranking.

Provider and offerWhat it says it coversWho tests, as statedRetest, as statedWorth knowing
Bishop Fox, AI/LLM Security AssessmentRunning apps and LLM endpoints, web and API, agent permissions; retrieval and agents by scope"Our consultants""Retesting to confirm that remediation efforts are effective"; no count or deadlineAlso covers model theft and training-data attacks
Blaze, AI penetration testing"LLM applications, RAG pipelines, and agents", plus web, API, identity and data accessNamed researchers; "manual validation of every finding""Fix validation when included in scope"A retest is not automatic. Ask for it in the quote
Bugcrowd, penetration testing for AI"Any LLM implementation or other AI use case"; checks the OWASP Top 10 for LLMsA matched team from its crowdNot stated for the AI testThe page lists little scope detail
Cobalt, AI & LLM PentestLLM-enabled apps, the networks hosting them, API connections and web appsCobalt Core testers; "over 3-dozen" with LLM experienceIts pricing table shows 6 or 12 months; its FAQ says unlimited during the contract. Ask which appliesSold as annual credits. One credit is "the equivalent of 8 hours" of testing, which is not a promise of eight human hours
HackerOne, AI Red Teaming"Prompts, models, APIs, and integrations"; retrieval pipelines and agent workflowsVetted researchers; the company says more than 750 focus on AIRetest and validation mentioned; terms not statedRuns as "15 or 30-day engagements". Its page cites the 2025 OWASP edition
NetSPI, AI/LLM penetration testingFive separate services, from LLM web app testing to custom model testingNetSPI security expertsNot stated"Continuous AI Findings Validation" checks the output of AI security tools. It does not test your LLM app
Qualysec, LLM penetration testingChatbots, copilots, document systems, agents, plugins, fine-tuned modelsNot stated"Remediation testing" is a listed step; no termsReport plus a letter of attestation; "typically takes 1–3 weeks". Its cost section gives no number
Rhymetec, LLM penetration testingRetrieval pipelines and vector databases, agent tool and plugin access, API integrations"Both automated and manual" testingNot stated on the LLM service pageAsk whether retesting is available and what it costs
Synack, AI and LLM PentestingAI and LLM apps, plus common web exploitsSynack's researcher networkNot statedNo price is listed for this service. Do not confuse it with Sara
Trail of Bits, AI/ML SecurityTraining data, pipelines, model files and deployed agent loopsIn-houseAskA deep review of models and infrastructure, more than a quick app test

We did not find an open LLM sample report among the pages we read. Cybersecify publishes a sample without a form, but it is a web and API report with no LLM findings in it. So ask every provider on your list for a redacted report from an LLM engagement before you sign. It is the fastest way to see whether you will get evidence your customer can follow.

What should you send providers before asking for a quote?

Send every provider the same description of your AI feature. Then each one prices the same job, and their answers line up.

Add these lines to your normal scope. Leave out passwords, keys and the text of your system prompt.

  1. What the feature does, in one sentence.
  2. Where users reach it: web app, API, chat tool, email.
  3. Model: which provider, and whether they host it or you do.
  4. Data it reads: which sources, and how one customer's data is kept from another's.
  5. Who can add content it reads: staff only, customers, or the public web.
  6. Roles and test customers: the user roles to test, and two made-up customers with made-up data for the cross-customer check.
  7. Tools and actions: every tool it can call and what each can change.
  8. Outside connectors: plugins, MCP servers, other agents, and who runs each.
  9. Where to test: production or a copy, and which actions must never really happen (refunds, emails, deletions).
  10. Cost and uptime attacks: in or out.
  11. What surrounds it: whether the web app and API around the feature are in this test or a separate one.
  12. Fix checking: how many retests you need, by what date, and whether you need an updated report.
  13. Who reads the report, what they asked for in their own words, and your deadline.
  14. Ask each provider to mark every line as covered, not covered or needs a conversation, with a price for anything extra.

A filled-in example (made up). "Support assistant inside our web app. Hosted model from a third party. Reads each customer's uploaded help documents. Staff and customers can add documents. Three roles: customer user, customer admin, our support staff. Two test customers with made-up data. Two read-only tools: order lookup and ticket lookup. Test on staging only. Cost attacks out. Web app and API were tested in March and are out of scope. One retest needed about 45 days after findings, with an updated report. Our customer's security team reads the report and asked for 'a penetration test of the AI feature'. Report needed in six weeks."

AI scope lines

  1. What the feature does:
  2. Where users reach it:
  3. Model:
  4. Data it reads:
  5. Who can add content it reads:
  6. Roles and test customers:
  7. Tools and actions:
  8. Outside connectors:
  9. Where to test:
  10. Cost and uptime attacks:
  11. What surrounds it:
  12. Fix checking:
  13. Who reads the report:
  14. Ask each provider to mark every line as covered, not covered or needs a conversation, with a price for anything extra.:

This list prepares a conversation with providers. It does not authorize testing.

Prepared with The PenTest Index: https://thepentestindex.com/llm-penetration-testing-services/

If you cannot answer a line yet, write "Not known yet, please advise." A blank reads as "none," and that is how gaps get missed.

This list starts a conversation. It does not give anyone permission to test. Written authorization has to name the real systems and the real actions, and it comes from you in the contract or the rules of engagement. If a third party hosts your model, check their terms on security testing before work starts.

Once quotes come back, compare them line by line. For the rest of your scope, see a filled-in penetration testing scope.

Is AI red teaming the same as LLM penetration testing?

Providers use the two names for overlapping work, so compare the scope and ignore the label. Schellman calls its service AI Red Teaming and prices it per engagement at about two weeks. HackerOne's AI Red Teaming runs for 15 or 30 days. Invadel calls similar application work a penetration test and treats model-level red teaming as extra scope, "quoted after scoping."

A rough way to tell them apart: an application test asks "can someone misuse this feature to reach data or actions they should not?" A model-level exercise asks "can this model be pushed into harmful behavior at all?" Most companies building on someone else's model need the first one.

OWASP publishes a free guide to judging these providers, and it is written to help you separate real adversarial testing from "jailbreak-only" offerings: OWASP Vendor Evaluation Criteria for AI Red Teaming Providers & Tooling (v1.0, February 2026). For the general difference, see red teaming vs pentesting.

Do you have to buy a separate LLM penetration test?

Not always. It depends on what the person asking for the report needs and what your current test already covers.

Provider pages name frameworks such as the EU AI Act, ISO/IEC 42001 and the NIST AI Risk Management Framework. A provider naming a framework does not mean the framework requires this test. Your customer, auditor or regulator decides what counts. No provider and no index can promise a report will be accepted.

So ask the person who asked you:

  • "Which system must the test cover, and does it include the AI feature specifically?"
  • "Is automated testing alone acceptable, or do you need people doing the testing?"
  • "Do you need proof that findings were fixed and rechecked, and by when?"

Their answers go straight into line 13 of your scope. If they say your existing web app test is enough once it includes the AI feature, you may only need an add-on from your current provider.

Can you just run a tool yourself?

Yes for regular checks between tests. No if someone outside your company needs an independent report.

Promptfoo, garak and PyRIT are open-source tools you run against your own feature. They are good at repeating known attacks every time you change a prompt or a model. They do not understand your business rules, they do not check whether one customer can reach another's documents unless you build that test, and nobody independent signs the result. Many teams use both: a tool in the release process and a commissioned test once or twice a year. The same split applies across security testing; see penetration testing vs vulnerability scanning.

More questions buyers ask

How long does an LLM penetration test take?

Stated times on the pages we read: Cybersecify says 5 business days per scope, Invadel about a week for one feature, Schellman around 2 weeks, Qualysec 1 to 3 weeks, and HackerOne 15 or 30 days. These are testing periods the providers state. They are not booked start dates. Add scheduling time, which Software Secured puts at 3 to 6 weeks, and ask for the report date in writing.

Do testers need our system prompt or source code?

It depends on the offer. Some tests work only from the outside, the way a user would. Others go faster and deeper with the prompt, the tool list and test accounts. Ask each provider what access their quoted price assumes, and share anything sensitive through the channel you agree with them, not in a scope document.

Is it safe to test an agent in production?

Only with written limits. Agree which actions must never really run, set rate and spending caps, and name someone who can stop the test. A staging copy with made-up customer data is the safer choice when your agent can send, change or delete things. Our guide to rules of engagement covers what that document should say.

Does our model provider's own testing cover us?

No. The model maker tests its model. It has not tested the documents your feature reads, the permissions behind your tools, or what your app does with the answer. Those parts are yours.

Is a provider missing, or has an offer changed?

Send us the provider name and the source page. We check corrections against the evidence and update the entry and its check date.

How we checked, and how we make money

We read each provider's own service and pricing pages on October 10, 2026 and recorded what the page says, including where two pages from the same company disagree. Everything in the tables is provider-published. We have not bought these services, inspected a real engagement or confirmed any term in writing with a provider. We do not perform or authorize penetration testing. Read our method.

Provider links on this page are plain links. Payment never decides which providers appear, the order or which offer fits. See how we make money.

Sources

SourceWhat we used it forChecked
Astra pricingPentest Expert price, target types, re-scansOctober 10, 2026 (re-scan request rule and Pentest Auto: October 7–8, 2026)
Bishop Fox AI/LLM Security AssessmentScope, retest wordingOctober 10, 2026
Blaze AI penetration testingScope, fix validation wordingOctober 10, 2026
Bugcrowd AI pen testScope and method wordingOctober 10, 2026
Cobalt AI & LLM Pentest, Autonomous Pentest, pricingScope, credit definition, retest statementsOctober 10, 2026
Cybersecify pricing and service pagePrices, scope rules, retest, exclusionsOctober 10, 2026
HackerOne AI Red TeamingScope, engagement lengthOctober 10, 2026
Invadel AI/LLM pricingTier prices and definitions, retest, timingOctober 10, 2026
NetSPI AI/ML penetration testingService listOctober 10, 2026
Qualysec LLM penetration testingScope, deliverables, durationOctober 10, 2026
Rhymetec LLM penetration testingScope, retest wordingOctober 10, 2026
Schellman AI Red TeamingMinimum price, duration, exclusionOctober 10, 2026
Software Secured pricing and AI pentestingStarting price, retest rounds, schedulingOctober 10, 2026
Synack pricing and AI and LLM PentestingSara asset type; LLM serviceOctober 10, 2026
Trail of Bits AI/ML SecurityService descriptionOctober 10, 2026
OWASP GenAI LLM Top 10 2026 and the 2026 listEdition and entry namesOctober 10, 2026
OWASP Vendor Evaluation Criteria for AI Red Teaming Providers & Tooling v1.0Linked as an official buyer resourceOctober 10, 2026