LlamaIndex Applications Penetration Testing
A LlamaIndex agent picks a tool from its name and docstring alone, not a permission the framework enforces. We test what each tool, retriever and shared context can actually reach. CREST-certified testers, fixed price from £6,000 for a 5-day single-application scope, quoted within 24 hours.
- Unlimited retesting
- Unlimited pre-retesting
- No hidden fees
“I would highly recommend EJN Labs to any organisation seeking reliable, detailed, and well-managed penetration testing services, particularly for government or enterprise-level projects.”
“There wasn’t another company we could find that could deliver what we needed in the timeframe we needed. The client loved it, and we got instant ROI from the engagement.”
LlamaIndex documents that an agent chooses which tool to call, and the arguments it writes, based on that tool’s name and docstring rather than any permission the framework enforces itself. We test what each tool can actually do once called, not what its description implies.
Why LlamaIndex security comes down to what a tool’s docstring is trusted to do
LlamaIndex’s retriever documentation describes a retriever as responsible for fetching the most relevant context for a query, and filtering that result down to one tenant’s own data is documented separately, per vector store, as a Metadata Filter guide for Pinecone, Qdrant, Weaviate, Milvus and Neo4j among others. LlamaIndex documents one integration, Nile, specifically as a multi-tenant PostgreSQL vector store, which is a strong signal that the rest were not built with tenant isolation as a default; on those, a metadata filter has to be applied correctly by your own code on every query. We test whether a query from one account can retrieve chunks that belong to another.
An agent’s choice of tool is documented, not enforced: LlamaIndex’s agent guide is explicit that a tool can be anything from an arbitrary Python function up to a full query engine, and that the model relies on the tool’s name and docstring, tuned by the developer, to decide when to call it. Pausing for approval is available through LlamaIndex’s human-in-the-loop pattern, built on InputRequiredEvent and HumanResponseEvent, but only for the specific tools a developer wires up that way, and everything else runs automatically. The underlying Workflow engine lets any step update shared state on its Context object, so data one step collected can be visible to a step that never should have needed it. We test what a tool actually reaches once called, and what a shared workflow Context is carrying between steps.
LlamaIndex applications increasingly reach outside the process entirely. Its LlamaCloud API key documentation confirms that Parse, Extract and Index are hosted services you call with an API key shown once at creation, so a document sent through LlamaParse leaves your infrastructure for LlamaCloud’s servers unless you run parsing yourself, and a leaked key cannot simply be viewed again to confirm what it was. LlamaIndex’s own workflows can also be exposed as MCP servers or call another MCP server’s tools, adding a remote tool-calling surface on top of the ones defined locally. It is the same tool-and-data-boundary question we test on our LangChain engagements, applied to LlamaIndex’s own retrievers, agents and cloud services.
SCOPE
What we pen test on a LlamaIndex application
Retriever and Query Engine Access Across Tenants
LlamaIndex’s retriever documentation describes a retriever as fetching the most relevant context for a query, with no tenant or user check built into that call by default. We test whether a query issued as one account can be made to retrieve chunks or nodes that belong to another.
Metadata Filtering Is Opt-In and Store-Specific
Each vector store integration LlamaIndex supports, Pinecone, Qdrant, Weaviate and Milvus among them, has its own Metadata Filter guide, and only one integration, Nile, is documented specifically as a multi-tenant PostgreSQL store. We test whether the metadata filter your application applies is actually present and correct on every retrieval path, not just the ones exercised during development, the same tenant-isolation question we test on our LangChain engagements.
Tool Selection Driven by Name and Docstring
LlamaIndex’s agent guide states plainly that tool selection and the arguments written for it rely strongly on the tool’s name and description, tuned by whoever built it, rather than a permission the framework checks. We test whether a tool’s actual access matches what its name and docstring imply, and what happens when a model is induced to call it in a way its description did not anticipate.
FunctionTool and ToolSpec Scope
A LlamaIndex Tool can be an arbitrary Python function wrapped by FunctionTool or a full query engine exposed as a tool, and LlamaIndex notes that the credential or connection behind that function is whatever the function itself was written to use. We test what a FunctionTool’s underlying function can reach, from a database connection to a file path, against what the agent’s task actually needs.
Human-in-the-Loop Is Opt-In Per Tool
LlamaIndex’s human-in-the-loop pattern pauses a workflow before a specific tool call using an InputRequiredEvent and a HumanResponseEvent, built by a developer around a named tool such as a dangerous task that requires confirmation. We test which tool calls in your agent are actually gated behind that pattern and which execute without any pause.
MCP Server and Tool Integration
LlamaIndex documents converting its own workflows into MCP servers and calling tools exposed by another MCP server from inside an agent, adding a remote tool-calling boundary on top of the ones defined in the application itself. We test what an MCP tool call can reach on both sides of that boundary.
Workflow Context Shared State Across Steps
LlamaIndex’s Workflow engine lets any step read or write shared state on a Context object and send events directly from it, which is how data collected in one step becomes available further down the workflow. We test whether a Context ends up carrying data into a step, or a branch, that should never have had access to it.
LlamaCloud API Keys for Parse, Extract and Index
LlamaCloud’s documentation is explicit that an API key is shown in full only once, at the moment it is generated, and cannot be viewed again after you leave that page, for Parse, Extract and Index alike. We test where each key is stored in your application and what a copy of it would let someone else do against your LlamaCloud project.
LlamaParse as a Hosted Document Pipeline
LlamaParse is documented as one of LlamaCloud’s hosted API services rather than a local library, so a document processed through it is sent to LlamaCloud’s servers for parsing unless you run an alternative parser yourself. We test what document types and fields actually pass through LlamaParse against what your data-handling policy allows leaving your infrastructure.
Prompt Injection into Tool-Calling Decisions
OWASP’s LLM Top 10 treats direct and indirect prompt injection as the route by which content in a user message or a retrieved document reaches an agent’s tool-calling decision, and LlamaIndex’s retrieval-first design means an indexed document is a routine source of that content. We test both paths, including injected instructions sitting inside a document your retriever returns.
OUR PROCESS
LlamaIndex Applications Penetration Testing: From Scope to Attestation
Scope and Access
We agree which application, model provider, vector store and tool set are in scope, plus credentials for every user role and any LlamaCloud API keys your indices or LlamaParse pipeline depend on.
Tool and Retrieval Mapping
We map every tool, FunctionTool and query engine an agent can call against what its underlying function or connection actually permits, and check whether retrieval calls and workflow Context state are scoped by user or tenant.
Manual Testing
A CREST-certified tester manually attempts prompt injection, tool misuse and metadata-filter bypass, chaining findings where a single weak boundary lets one issue reach another.
Attestation and Retest
You get a technical report with CVSS scores and reproduction steps, a walkthrough call, a free retest once fixes are deployed, and an attestation letter for auditors.
CREDENTIALS
Verified Accreditations Auditors Accept
Every credential below is independently verifiable. UK procurement teams, FCA supervisors, ISO 27001 / SOC 2 auditors, and cyber insurance underwriters all recognise these standards.
GET YOUR QUOTE
Get a CREST LlamaIndex pen test quote in 24 hours
A fixed-price quote back in one business day, from a named CREST assessor. No sales pipeline, no chasing.
- CREST and IASME accredited. Testing your auditors and clients already recognise.
- Fast-track testing within 24 hours where required. Free retest of every fix included.
- Live findings via your client portal, not a four-week PDF.
- Fixed price from £3,500 for a single-role, single-app scope, agreed up front. Most engagements run £5,000 and up. No day-rate surprises.
Under NDA Further named references available on a scoping call.
- We reply within one business day with a fixed-price quote from a named CREST assessor.
- You approve the scope and we book a start date, usually within 24 hours.
- Live findings land in your client portal as we test, with a free retest of every fix.
Get your fixed pen test quote in 24 hours
Quote request received
We will reply within one business day with your fixed-price quote from a named CREST assessor.
Your data stays with us. No newsletter signup.
or book a 20-min scoping call first
We reply within one business day. Your data stays with us. No newsletter signup.
COMPLIANCE READY
Reports Mapped to Every Framework
Findings are written so your team can reference the report against each framework without translation work.
ISO 27001:2022
Annex A.8.8 management of technical vulnerabilities plus A.5.15-5.18 and A.8.2-8.5 access control validation.
SOC 2 Type I & II
CC6 logical access, CC7 system operations, CC8 change management evidence.
PCI DSS
Requirement 11.4 application penetration testing across cardholder data environments, including ecommerce penetration testing for online retail platforms.
FCA SYSC
SYSC 4.1.1R, 6.1.1R, 13 mapped to each finding for FCA-regulated firms.
UK GDPR
Article 32 effectiveness testing, customer-data security controls, ICO-acceptable evidence.
Cyber Essentials Plus
Direct certification through our IASME body status, single-vendor delivery.
PRICING
Transparent LlamaIndex Applications Penetration Testing Pricing
Pricing depends on the number of roles, integrations and environments in scope. See our pricing page for how we quote.
Depends on AI system complexity
Single LLM-powered chatbot, basic RAG (≤100 documents), no agent tools. Around 5 to 7 working days from kickoff to report.
Get a fixed quoteDepends on AI system complexity
Multi-tool agent system, complex RAG pipeline, fine-tuned model, multi-tenant. Around 8 to 12 working days from kickoff to report.
Get a fixed quoteDepends on AI system complexity
Production AI platform, multi-agent orchestration, regulated AI use case (FCA, NHS), custom-trained models. Around 12 to 18 working days from kickoff to report.
Get a fixed quoteSECTORS
Sectors We Test LlamaIndex For
Sector-specific scoping for regulated UK organisations.
Fintech & FCA-Regulated
FCA SYSC, Open Banking FAPI 1.0, PSD2 SCA, payment-flow scrutiny, KYC/AML testing.
Fintech sector pageSaaS Companies
SOC 2 Type I & II evidence, multi-tenant boundaries, role escalation, customer-tenant isolation.
SaaS sector pageLaw Firms
SRA Cyber Standard, privileged data, conveyancing fraud defence, partner-tier procurement.
Law firm sector pageHealthcare
NHS DTAC, DSP Toolkit v6, UK GDPR Article 32, EHR systems, telehealth platforms.
Healthcare sector pageInsurance
FCA / PRA Operational Resilience, cyber underwriting, claims data, broker portals.
Insurance sector pagePublic Sector
CCS / G-Cloud framework, NCSC-aligned, citizen-facing services, PSN-compliance scrutiny.
Public sector pageWHY EJN LABS
What You Get From LlamaIndex Applications Penetration Testing
Six concrete differentiators competitors don’t all match.
CREST-Certified Testers, Verifiable
Every test by a CREST-certified pen tester (CRT, CCT APP, CCT INF where applicable). Verify our company status at crest-approved.org.
24-Hour Startup, Where Required
From signed scope to active testing in a single business day for incident response, audit deadlines, or regulator-driven timelines.
Live Findings, Not 4-Week PDFs
Critical issues reported during testing through your client portal. Your team remediates while testing continues.
Audit-Ready Reports
Executive summary plus full technical report with CVSS scores and explicit framework mappings (ISO 27001, SOC 2, PCI DSS, FCA SYSC).
Free Retests, Standard
Verify remediation of every finding before close-out. Letter of attestation for audit submission included. Most competitors charge £1,500-£3,000 per retest.
UK-Based CREST Testers
Every engagement performed by vetted, UK-based CREST-certified testers, matched to your needs, security clearance, and compliance scope.
FAQ
Frequently Asked
What access do you need to test our LlamaIndex application?
We need working credentials for at least one account per role your application exposes, plus visibility into the tools and query engines an agent can call, including which ones hold a write-capable database connection or a file-system path. If you use LlamaCloud for Parse, Extract or Index, we also need to know which pipelines are in scope.
Will testing touch our live data?
We test whichever environment you give us access to. If that is production, we agree exclusions upfront, such as destructive tool calls, real document uploads to LlamaParse, or outbound emails, and we do not run untested prompts against real customer records or a production vector store without that agreement in writing.
How long does a LlamaIndex application penetration test take?
A single LlamaIndex application sits in our 5-day single-application scope, with a report typically landing around 5 to 7 working days after kickoff. An application with a multi-tool agent, multiple tenants or a fine-tuned model moves into a wider AI penetration testing scope with more testing days.
Do you test LlamaIndex running on our own infrastructure as well as LlamaCloud?
Yes. We test LlamaIndex agents, indices and workflows you run yourself the same way as an application built on LlamaCloud’s Parse, Extract and Index services. The tool, retrieval and workflow testing is the same either way; only the infrastructure-level checks around the hosting platform change, and those are scoped as cloud penetration testing where needed.
What is out of scope for a single-application LlamaIndex test?
The underlying model provider’s own infrastructure, whether that is OpenAI, Anthropic, Google or a self-hosted model, is out of scope; we test how your application uses the model, not the provider’s platform. A separate agent, MCP server or tool server that sits outside this application’s LlamaIndex integration is scoped and quoted separately.
Do you need our source code?
No. Testing is black-box against the running application by default. A grey-box option, where we review FunctionTool definitions, metadata filter configuration and Workflow Context usage, is available if you want faster or deeper coverage of specific findings.
Does LlamaIndex have a customer penetration-testing policy we need to follow?
LlamaIndex is an open-source library you deploy and control yourself, so there is no vendor notification process for the library itself. If your application uses LlamaCloud services such as Parse, Extract or Index, or a hosted model provider, we confirm that vendor’s current penetration-testing and acceptable-use terms during scoping before any testing begins.
Are your testers CREST certified?
Yes. Every LlamaIndex engagement is carried out by UK-based, CREST-certified testers, and your report and attestation letter are recognised by auditors and insurers accordingly.
20+ CREST-accredited testing services in one place
Web, mobile, API, cloud, AI, infrastructure, red team. Pick the test that fits your environment.
Get a fixed price for your LlamaIndex application
A LlamaIndex agent picks a tool from its name and docstring alone, not a permission the framework enforces. We test what each tool, retriever and shared context can actually reach. CREST-certified testers, fixed price from £6,000 for a 5-day single-application scope, quoted within 24 hours.



