The individual pieces of what I've been documenting aren't new. Confident fabrication, overclaiming, self-verification failing to catch its own errors — these are studied, published, and well understood. Alignment researchers and computer scientists have absolutely seen the underlying phenomena before.
What hasn't been done is putting the pieces together.
Rather than assume that, I checked it directly — ran two of the most frequent patterns in my tracker against the published literature, at the assembled level, not just the ingredient level.
Pattern one (highest instance count): 49 separate queries, searching for this exact pattern — as a single, combined behavior — described anywhere in the literature. No match found.
Pattern two (second-highest instance count): 25 queries. Here, the ingredients overlap substantially with an already-named academic category — "intrinsic hallucination" and "context-conflicting hallucination" (Ji et al. 2023; Zhang et al. 2023; Huang et al. 2023). But the assembled version — this specific pattern, occurring this way, in this sequence — still has no exact match.
The result cuts against the obvious explanation. If frequency drove discoverability, the most common pattern should be the easiest to find in the literature. It wasn't. It had zero matches. The second-most-common pattern got close at the ingredient level and still came up empty at the assembled level.
So: frequency inside sustained use doesn't predict whether something's already been named outside it.
Benchmark research is built for breadth — many short, independent prompts, aggregated into a rate. It isn't built to hold one person's ongoing, adversarial engagement with the same system over months, cross-referencing every claim against a running record. That kind of depth on a single, continuous case isn't how the field currently gathers evidence — so the assembled pattern doesn't show up in the literature, even though every piece of it does.
Not a claim to have discovered a new failure mechanism. The mechanisms are already known and cited. What's being contributed here is the assembled form — individually confirmed, tracked over time, checked against the literature at the level that actually matters: not "does this ingredient exist somewhere," but "has this exact combination been documented as what it is."
It hasn't. That's the gap this series is closing.