Trustli
Short version: we don't. Nothing installs, nothing integrates, and no software of mine ever runs inside your systems.
The whole value of an independent assessment is that the assessor stays outside. You hand me evidence; I never reach in and take it.
An endpoint I can talk to, with test credentials — the same instance you'd hand a QA tester. Not production. No real customer data.
My probe agents hold conversations with it and try to push it off-script: prompt injection in several styles, broken and failing tool calls, ambiguous "are you finished?" traps. The probes run on different underlying models from each other, because different models miss different things.
System prompt, tool and function definitions, the scopes and permissions it runs with, plus any spend or rate ceilings. A file, a repo link, or a redacted export — whatever is easiest.
Whatever your agent already writes. An export is fine; if you'd rather, a read-only viewer key to your logging tool works too. Redact freely — I'm looking at behaviour patterns, not content.
I run the fifteen checks. Where the evidence answers a check, it's graded. Where it doesn't, I come back with specific questions — never a generic questionnaire, always "your config shows the agent can do X, is there a ceiling on that, and where is it enforced?" That conversation takes about ten minutes and it's the only meeting required.
You get a signed verdict listing every check with its evidence, and a public badge page your customers' security teams can read in minutes. If something fails, you get the finding and what "fixed" looks like — free, and re-verification after you fix it is included, never billed.