Public record

Validation summary

Figures as of July 2026. Regenerated as the corpus grows. Written for skeptical readers — reporters, studio research departments, and buyers — with limitations stated alongside results.

The claim

An AI that watches the film and predicts real-world reception — with no human test audience.

aiScreeningRoom predicts how real-world audiences will receive an unreleased film by watching the actual footage — the entire film — with no human test audience.

That claim rests on a two-link chain, each link independently measured, neither ever fit to the other:

  1. iScreeningRoom panels predict public reception. Taste-matched human test audiences have a measured track record against the public ratings their films eventually earned.
  2. aiScreeningRoom predicts panel results — validated blind. The AI is tested on films it has never seen and could not recognize, against what real panels said about the same films.

No step in the chain has access to the answer it is being graded on.

The corpus

More than ten years of real test screenings.

35,666
Audience survey responses
42,936
Timestamped moment-by-moment reactions

Panels are nationwide audiences selected for interest in the film’s genre, premise, and comparable titles — a taste-matched sample of the film’s natural audience, not whoever showed up to a theater on one night.

Panel testing is ongoing; the corpus and every figure in this document grow with it.

Link 1

Panels predict public reception

For every corpus film that went on to earn an organic public IMDb rating — at least 100 votes, identity verified by exact ID (including retitled releases), and a ballot histogram free of organized voting — we compared the panel score against the rating the public eventually gave.

Across the 37 qualifying films:

97%
36 of 37 within 1.0 IMDb point
0.48
Mean absolute error (points)
ρ 0.62
Rank correlation vs. IMDb

20 of 37 landed within half a point. Under a stricter polarization screen, rank correlation rises to ρ = 0.69.

Why the polarization screen exists

A single IMDb mean can hide review bombing and fan stuffing. We record the share of ballots at 1 star and at 10 stars and exclude films where either pole holds 30% or more of all votes. Worked example: one politically charged faith film shows IMDb 1.6 from a histogram that is 71% one-star votes, against a panel score of 4.42 — that is protest voting, not audience reception. Without the screen, correlation drops to 0.40 — and the misses are dominated by films with organized-voting histograms, not by panel error. Public ratings are a noisy, manipulable proxy; a several-hundred-person taste-matched panel is a cleaner read of the same construct.

Link 2

The AI predicts panel results, validated blind

The condition that matters to a buyer is an unreleased film the model has never encountered. So the headline accuracy figure comes from the honest subset: films the model did not recognize, plus fully blind runs (title and metadata withheld).

On that subset (n = 10):

ρ 0.89
Rank correlation with panel
0.44
MAE on the panel’s 1–5 scale
7 / 10
Within half a star of the panel

8 of 10 within 0.75. The remaining two misses were underpredictions of about three-quarters of a star to one star — the known conservative bias on stronger films, not scatter in both directions. Including the misses, mean absolute error was 0.44.

The recognition control

If accuracy depended on the model recognizing famous films, it would be memory, not measurement. It doesn’t: the unrecognized-plus-blind subset is at least as accurate as the recognized set (ρ 0.89 vs 0.71). The model is judging footage.

Why the chain, not a direct AI→IMDb table

Only 7 AI-scored films have organic public ratings, all recognized — unreleased films don’t have public data, which is the entire reason a prediction is needed. The prospect-facing claim therefore composes two independently measured links: AI→panel (blind-validated) and panel→public (bridge study above). Nothing is fit to public data at any point.

Operating product

What production runs today — and what Link 2 still measures.

Production runs the v2.3t viewing-pass prompt (band-commitment scoring, promoted after a pre-registered holdout close-out in July 2026). The Link 2 figures above come from the earlier blind validation study (honest subset, n = 10) — they measure AI→panel accuracy under the honest condition, not a re-score of those ten films on the current prompt.

Holdout close-out asked a different question: whether the v2.3t fix generalized to films the system had never trained on. On eight holdout films (fresh v2.2.3 baseline ×3 vs v2.3t ×3), the pre-registered criteria confirmed: mid-band paired improvement and top-proxy relief without mid-band gate failures. That read spent the licensed holdout look for this prompt generation. Absolute accuracy claims for buyers continue to rest on Link 1 and Link 2 above; holdout confirms the operating prompt’s band-level behavior relative to its predecessor.

Methodology

Safeguards

  • The holdout wall is absolute. Holdout films are never trained on, never prompt-engineered from, never used to tune anything. They answer one question: does this work on films the system never learned from? One licensed look per candidate generation; peek-then-adjust is prohibited.
  • Public ratings never enter calibration. The panel mean is the sole calibration target; public ratings are validation only.
  • Every ground truth is construct-checked. Polarization screening, verified title identity, exclusion of client-invited audiences and trailer-only tests. No number is trusted because it’s official — only because its measurement survives inspection.
  • Honest claims come from honest subsets. The advertised accuracy figure is the unrecognized-plus-blind condition, because that is the buyer’s actual condition. If accuracy held only on recognized films, we would call it what it would be: contamination, not capability.
Known limitations

Stated here because a validation record that hides its weaknesses isn’t one.

  • The honest subset is small (n = 10) and growing. The free beta exists in part to grow it: predictions are logged before any panel or public outcome exists, then scored when those outcomes arrive. If accuracy regresses materially as n grows, the claims will be repositioned accordingly.
  • Top-band compression is reduced, not gone. Underpredicting films audiences love was the priority systematic error. The v2.3t prompt materially relieved it on train and on holdout (secular top-proxy holdout film: paired error improvement of 0.28 panel points). Residual compression and occasional mid-film overshoot remain possible; the defect is no longer the unaddressed ceiling it was before that cycle.
  • Bottom-band overprediction is holdout-confirmed. On weaker films, predictions skew high — train bottom-band bias rose under v2.3t, and holdout replicated the upward pressure (all four bottom-band holdout films moved up; candidate bias +0.233 on that set). A disclosed reporting transform is under pre-registered review and is not shipped; until and unless it ships with a train-fit / prospectively-unconfirmed disclosure, displayed scores at the low end carry this known skew. Raw scores remain the validation canonical.
  • Faith dramas and documentaries: public ratings are the unreliable side. In that segment, IMDb data is frequently distorted by culture-war voting or too thin to use. We report the panel-scale prediction and say why, rather than manufacture an IMDb forecast.
  • Perception inversions are a residual uncertainty class — partly teachable, not a hard ceiling. Some films succeed or fail on taste utilities a craft-and-engagement read underweights (genre appetite, matched-audience payoff, what “works” for this audience). Panel data — aspect scores, moment reactions, segment structure — is how those utilities get taught into the instrument (rubric weights, train-only exemplars, segment handling). That path has already shrunk some misses of that class without eliminating them. What remains after average audience utilities are learned defines the uncertainty band; reports disclose it rather than promising an oracle of taste.
How the record grows

Prospective ledger. Frozen at log time. Never rewritten.

Every beta film is eligible for a prospective accuracy ledger: prediction logged while no panel actual exists, frozen at log time, scored only when panel results later arrive — never rewritten. The ledger starts empty by design — existing corpus evaluations are refused as non-prospective. It fills as real client films are evaluated and later panel-tested or released.

Incumbent testing methods assert authority and publish no accuracy record. This document is ours, and it will be updated — including when the numbers move against us.

The accuracy record is the company. Integrity is the product.

Methodology questions: iscreeningroom.com