AI Surfaces in the real world
An AI surface can appear helpful while the application behind it exposes data or takes actions its user was never allowed to perform.
Test the whole application—not only the model.
The problem
The model is only one part of the application.
Your AI-enabled application may be able to expose sensitive data or take actions its user is not allowed to perform.
The real boundary is formed by identity, data, permissions, and tools. A polite refusal or detector warning does not tell you what the application attempted—or what happened downstream.
What can someone persuade your AI surface to see, call, change, or disclose?
The solution
AI Application Security Assessment
Test one authorized AI surface as an application—not just as a model.
The fixed-scope assessment follows meaningful attempts through identities, tools, services, and downstream effects. It identifies the owning control, helps the team correct the gap, and repeats the same meaningful test.
See how the assessment worksCase study
One HR chatbot. Four different security outcomes.
A synthetic HR chatbot used a service identity with more authority than its user. The test separated what the model said from what the application attempted, enforced, and changed downstream.
- Vulnerable baselineAn unsafe email side effect was visible.
- Detected without preventionA warning appeared, but the effect remained possible.
- Optionally blockedOne pinned attempt was stopped by an added control.
- RetestedThe authoritative readback showed no unsafe email side effect.
This bounded synthetic example does not prove that every prompt injection is blocked or establish universal safety.
Read the evidence and limitationsThe wider practice
Security is one part of operating AI well.
AI applications must also be affordable, buildable, and observable. VOF connects those concerns through four working flagships.
-
Security
Injectionator
AI application-security platform — test application behaviour and optionally defend live paths.
Read more → -
Finance
Stater
A framework for AI token economics — cost vs yield per unit of work.
Read more → -
Development
Coherence
Specification-to-implementation alignment for LLM-assisted development.
Read more → -
Operations
GenAIOps
Observability, incident response, and runbook patterns for AI applications.
Read more →
How engagements run
Assess. Plan. Execute.
One pillar at a time, in three shapes. Examples below; the full matrix lives on the Work page.
-
Assess
Healthcheck — fixed fee, 1 week
OWASP LLM Top 10 audit against your stack.
-
Plan
Roadmap — fixed fee, 2–3 weeks
Threat model and remediation roadmap.
-
Execute
Project — day rate, 4–12 weeks
Deploy guardrails, instrument apps, train the team.
Latest writing
From the field
-
Before and after AI Application Security testing: one HR chatbot
A bounded synthetic example of why model output, application enforcement, and downstream effect must be tested separately.
-
JUNIE: The Customer Proxy in Our Development Workflow
Why we are putting a customer-proxy agent inside pull requests, and why passing tests is no longer enough proof.
-
/handoff and /pickup: An Ephemeral Handover Doc Instead of Yet Another Doc To Rot
We borrowed /handoff from Matt Pocock, added /pickup, and now we have a paired Claude Code skill that gives us a deliberate, throwaway context bridge across /clear — instead of the usual cargo-cult HANDOFF.md that nobody updates and nobody trusts.
Start with one AI feature.
Tell us what the surface can see or do. We’ll identify a bounded first assessment.
Discuss one AI featureNeed the wider picture?
Assessment is one part of the practice. See the full Assess, Plan, Execute matrix across Security, Finance, Development, and Operations.
Explore the work