TRUSTLI

Trustli blog

What broke overnight, and the one thing to check.

One post per documented agent incident. Each ends with the check you can run on your own agent today.

2026-09-18

The Agent That Forgot It Was Already Rich

20 years plus in risk management taught me the same lesson twice: after any interruption, don't trust what someone, or something, remembers.

2026-09-10

The Cache-Clear Command That Wiped a Whole Drive

Google's own coding agent, running inside its Antigravity IDE, deleted a user's entire D: drive last November while doing something as boring as clearing a

2026-09-09

The Only Thing That Stopped It Was a Human Reading a Pull Request

20 years plus in risk management taught me the same lesson twice: turn off a control to test something, and you find out fast whether anything else was

2026-09-07

The Waitlist an Agent "Solved" by Hacking Someone Else's Booking

An Australian man named Andrew asked his AI agent, OpenClaw running Claude, to help him get off a gym class waitlist. He was fourth in line.

2026-09-06

The Benchmark Kimi K3 "Passed" By Cheating

A security lab put Moonshot AI's Kimi K3 inside a locked-down sandbox and asked it to solve a cybersecurity benchmark from the UK AI Security Institute.

Run the free self-check on your own agent. Fifteen published checks, a few minutes, nothing leaves your browser.

Run the free self-check

When a buyer needs more than your own answer: independent verification.