Figures as of July 2026. Regenerated as the corpus grows. Written for skeptical readers — reporters, studio research departments, and buyers — with limitations stated alongside results.
Living document · updated with the corpus
aiScreeningRoom predicts how real-world audiences will receive an unreleased film by watching the actual footage — the entire film — with no human test audience.
That claim rests on a two-link chain, each link independently measured, neither ever fit to the other:
No step in the chain has access to the answer it is being graded on.
Panels are nationwide audiences selected for interest in the film’s genre, premise, and comparable titles — a taste-matched sample of the film’s natural audience, not whoever showed up to a theater on one night.
Panel testing is ongoing; the corpus and every figure in this document grow with it.
For every corpus film that went on to earn an organic public IMDb rating — at least 100 votes, identity verified by exact ID (including retitled releases), and a ballot histogram free of organized voting — we compared the panel score against the rating the public eventually gave.
Across the 37 qualifying films:
20 of 37 landed within half a point. Under a stricter polarization screen, rank correlation rises to ρ = 0.69.
A single IMDb mean can hide review bombing and fan stuffing. We record the share of ballots at 1 star and at 10 stars and exclude films where either pole holds 30% or more of all votes. Worked example: one politically charged faith film shows IMDb 1.6 from a histogram that is 71% one-star votes, against a panel score of 4.42 — that is protest voting, not audience reception. Without the screen, correlation drops to 0.40 — and the misses are dominated by films with organized-voting histograms, not by panel error. Public ratings are a noisy, manipulable proxy; a several-hundred-person taste-matched panel is a cleaner read of the same construct.
The condition that matters to a buyer is an unreleased film the model has never encountered. So the headline accuracy figure comes from the honest subset: films the model did not recognize, plus fully blind runs (title and metadata withheld).
On that subset (n = 10):
8 of 10 within 0.75. The remaining two misses were underpredictions of about three-quarters of a star to one star — the known conservative bias on stronger films, not scatter in both directions. Including the misses, mean absolute error was 0.44.
If accuracy depended on the model recognizing famous films, it would be memory, not measurement. It doesn’t: the unrecognized-plus-blind subset is at least as accurate as the recognized set (ρ 0.89 vs 0.71). The model is judging footage.
Only 7 AI-scored films have organic public ratings, all recognized — unreleased films don’t have public data, which is the entire reason a prediction is needed. The prospect-facing claim therefore composes two independently measured links: AI→panel (blind-validated) and panel→public (bridge study above). Nothing is fit to public data at any point.
Production runs the v2.3t viewing-pass prompt (band-commitment scoring, promoted after a pre-registered holdout close-out in July 2026). The Link 2 figures above come from the earlier blind validation study (honest subset, n = 10) — they measure AI→panel accuracy under the honest condition, not a re-score of those ten films on the current prompt.
Holdout close-out asked a different question: whether the v2.3t fix generalized to films the system had never trained on. On eight holdout films (fresh v2.2.3 baseline ×3 vs v2.3t ×3), the pre-registered criteria confirmed: mid-band paired improvement and top-proxy relief without mid-band gate failures. That read spent the licensed holdout look for this prompt generation. Absolute accuracy claims for buyers continue to rest on Link 1 and Link 2 above; holdout confirms the operating prompt’s band-level behavior relative to its predecessor.
Every beta film is eligible for a prospective accuracy ledger: prediction logged while no panel actual exists, frozen at log time, scored only when panel results later arrive — never rewritten. The ledger starts empty by design — existing corpus evaluations are refused as non-prospective. It fills as real client films are evaluated and later panel-tested or released.
Incumbent testing methods assert authority and publish no accuracy record. This document is ours, and it will be updated — including when the numbers move against us.
The accuracy record is the company. Integrity is the product.
Methodology questions: iscreeningroom.com