ChronoRealm is a local workspace being developed to help people check AI-assisted work: Where did this claim come from? What changed? What still needs checking?
Open a claim → inspect its evidence → record a correction → review the history.
For example, imagine an AI-assisted report says, “All checks passed.” A reviewer examines the supporting report and finds that some checks were skipped. The reviewer records a narrower conclusion: “The executed checks passed; the skipped checks remain untested.” The original claim, supporting evidence, and correction remain available for review.
This is an illustrative example of the intended workflow, not a reported human-pilot outcome. People still need to judge whether the evidence is reliable and sufficient.
Teams using AI for research need to check whether a conclusion is actually supported and whether later corrections have been accounted for. My hypothesis is that keeping claims, evidence, and corrections together could make that review easier. The benefit needs to be measured against ordinary document review. A more elaborate record could also slow reviewers down or create false confidence in poor evidence; the pilot should look for those failures.
I am Austin Simpkins, an independent researcher and developer. I am seeking $10,000 for a focused validation and pilot milestone, with full-time work on the funded pilot and a proposed eight-week schedule after funding and scope agreement.
The proposed budget is $4,000 engineering and packaging; $2,500 scoped external review; $1,500 infrastructure; $1,000 pilot work and documentation; and $1,000 contingency. The external-review allowance is a planning estimate, not a vendor commitment.
The proposed evaluation would involve 3–5 consenting reviewers and examine unsupported-claim detection, citation errors, missing authorization, review time, and correction-history usability. This small study would be exploratory. Independent evaluation and customer demand remain open questions.
The September 11, 2026 internal development report records 39/39 suites passing and 2,596 tests counted. Those totals include overlapping checks and skipped tests. These results describe bounded internal development checks; broader acceptance testing and independent review remain outstanding.
I welcome sponsors, technical reviewers, pilot partners, and criticism of the evaluation design. What would count as a meaningful improvement over existing document workflows? Which failure cases should be included?
Read or support the Manifund proposal
Public funding brief and test summaries
Writing disclosure: Prepared with AI assistance from project information supplied by Austin Simpkins.