Public record

Validation summary

Figures as of August 2026. Regenerated as the corpus grows. Written for filmmakers and skeptical readers alike - plain meaning first, every number kept, limitations stated alongside results.

The claim

An AI that watches the film and predicts real-world reception - with no human test audience.

aiScreeningRoom predicts how real-world audiences will receive an unreleased film by watching the actual footage - the entire film - with no human test audience.

That claim rests on two independently checked links. Neither ever peeks at the answer it is graded on:

  1. Do taste-matched panels predict public IMDb? Human test audiences selected for the film’s natural audience have a measured track record against the ratings those films eventually earned in the wild.
  2. Does the AI predict those panel results? Graded on films locked away from training and tuning - never trained or tuned on.

No step in the chain has access to the answer it is being graded on.

The corpus

More than ten years of real test screenings.

35,666
Audience survey responses
42,936
Timestamped moment-by-moment reactions

Panels are nationwide audiences selected for interest in the film’s genre, premise, and comparable titles - a taste-matched sample of the film’s natural audience, not whoever showed up to a theater on one night.

Panel testing is ongoing; the corpus and every figure in this document grow with it.

Link 1

Do taste-matched panels predict public IMDb?

For every corpus film that went on to earn an honest organic public IMDb rating - at least 100 votes, identity verified by exact ID (including retitled releases), a ballot histogram free of organized voting, and a panel of at least 100 respondents - we compared the panel score against the rating the public eventually gave. Each film is graded by a mapping that never saw it.

Across the 36 qualifying films:

89%
32 of 36 within 1.0 IMDb point
0.49
Average miss (IMDb points)
ρ 0.63
When panels ranked high, IMDb usually did

27 of 36 (75%) within three-quarters of a point; 23 of 36 (64%) landed within half a point. Under a stricter bombing screen, that ranking agreement rises to ρ = 0.70. No film is ever graded by a mapping that saw it, and the mapping preserves the scale’s spread - it can call a film exceptional or a disaster rather than compressing every prediction toward the middle.

Why we throw out some IMDb pages

A single IMDb mean can hide review bombing and fan stuffing. We record the share of ballots at 1 star and at 10 stars and exclude films where either pole holds 30% or more of all votes. Worked example: one politically charged faith film shows IMDb 1.6 from a histogram that is 71% one-star votes, against a panel score of 4.42 - that is protest voting, not audience reception. Without the screen, correlation drops to 0.40 - and the misses are dominated by films with organized-voting histograms, not by panel error. Public ratings are a noisy, manipulable proxy; a several-hundred-person taste-matched panel is a cleaner read of the same construct. The screen keys on a single pole. Some pages that pass it still show a U-shape - elevated shares at both one star and ten stars - so the published mean is pulled by the poles while the middle of the ballot (2-9) is closer to ordinary taste. On the same 36 films, using that middle-of-ballot mean only for titles meeting a pre-stated U-shape rule, 34 of 36 (94%) land within one point.

Manipulation weighting

IMDb’s own fraud detection visibly discounts distrusted ballots - the spread between a film’s raw ballot mean and its published rating. Films carrying large discounts have their influence on the mapping scaled down in proportion to that evidence (no film is deleted by it). The screen keys only on IMDb-internal evidence, never on disagreement with the panel.

Link 2

Does the AI predict panel results?

On 9 sealed holdout films - locked away from training and tuning; never trained on, never tuned on - graded against real panel results (3 viewing replicates per film, average reported):

ρ 0.750
When the AI ranked higher, panels usually did
0.285
Average miss on the panel’s 1–5 scale
9
Films locked away, ×3 watches each

Slightly conservative on average (bias −0.116) - the misses ran low, not random.

Cheat-check - is it remembering titles?

A skeptic will ask: is this just recognizing famous movies? We tested that. We asked, with no footage, what the system already knew about a film’s public reception. On 44 films it could not recall, it was about as accurate as on films it could recall (off by 0.26 vs 0.28 on the 1 to 5 scale). So the score is coming from the picture, not from the title.

Those 44 still include films used while we built the method. The accuracy number is the 9 films locked away from that work. That is why 0.285 on 9 films is the claim, not 0.26 on 44.

Why the chain, not a direct AI→IMDb table

Unreleased films don’t have public data - which is the entire reason a prediction is needed. The prospect-facing claim therefore composes two independently measured links: AI→panel (sealed holdout) and panel→public (bridge study above). The AI is never calibrated on public ratings - and each bridge film is graded by a mapping that never saw it.

Operating product

What production runs - and what Link 2 measures.

Production runs the v2.3t viewing-pass prompt (band-commitment scoring). The 9-film figures above (average miss 0.285, ρ 0.750) were produced under that v2.3t prompt. Buyer-facing accuracy claims rest on Link 1 and Link 2 above - the sealed holdout for AI→panel, not a re-score of films used to build the method. The licensed holdout look for this operating prompt is spent; further method work requires a newly licensed look.

Methodology

Safeguards

  • The holdout wall is absolute. Films locked away are never trained on, never tuned on. They answer one question: does this work on films the system never learned from? One licensed look per method change; peek-then-adjust is prohibited.
  • Public ratings never enter calibration. The panel mean is the sole calibration target; public ratings are validation only.
  • Every ground truth is construct-checked. Polarization screening, verified title identity, exclusion of client-invited audiences and trailer-only tests. No number is trusted because it’s official - only because its measurement survives inspection.
  • Honest claims come from honest conditions. Accuracy figures are only claimed under conditions that rule out the shortcut: films locked away for the AI, each bridge film graded by a mapping that never saw it. The cheat-check above is not a second accuracy claim. 0.285 on 9 films is the claim.
Known limitations

Stated here because a validation record that hides its weaknesses isn’t one.

  • The AI→panel accuracy record is still early (n = 9 films locked away). The free beta exists in part to grow it: predictions are logged before any panel or public outcome exists, then scored when those outcomes arrive. If accuracy regresses materially as n grows, the claims will be repositioned accordingly.
  • Public ratings for small films are manipulable ground truth. Paid vote-lifting and brigading exist; IMDb’s weighting corrects what it can detect, and our screens and weighting handle the visible cases. The projection assumes an organically rated release, and films positioned or marketed beyond their natural audience (e.g. marketing a drama as a comedy) tend to land below the projected center.
  • Films audiences love can be under-called. Top-band compression and occasional mid-film overshoot are possible.
  • Bottom-band overprediction is holdout-confirmed. On weaker films, predictions skew high - all four bottom-band holdout films moved up (candidate bias +0.233).
  • Faith dramas and documentaries: public ratings are the unreliable side. In that segment, IMDb data is frequently distorted by culture-war voting or too thin to use. We report the panel-scale prediction and say why, rather than manufacture an IMDb forecast.
  • Perception inversions are an uncertainty class - partly teachable, not a hard ceiling. Some films succeed or fail on taste utilities a craft-and-engagement read underweights (genre appetite, matched-audience payoff, what “works” for this audience). Panel data - aspect scores, moment reactions, segment structure - is how those utilities get taught into the instrument (rubric weights, train-only exemplars, segment handling). Misses of that class define the disclosed uncertainty band; reports disclose it rather than promising an oracle of taste.
How the record grows

Prospective ledger. Frozen at log time. Never rewritten.

Every beta film is eligible for a prospective accuracy ledger: prediction logged while no panel actual exists, frozen at log time, scored only when panel results later arrive - never rewritten. The sealed 9-film figures above stay as published. When a beta film later has panel results, that row is scored on the prospective ledger. We will report that larger set as a separate living figure. It does not rewrite the sealed 9. Existing corpus evaluations are refused as non-prospective. Your title is never published. No film, report, score, or excerpt is identified publicly without your prior written consent. Published figures are aggregate.

When a published figure on this page changes, the prior number stays in the log below.

  • 10 July 2026. 8 films locked away. Average miss 0.286. Rank agreement 0.762.
  • 26 August 2026. A ninth film’s panel score arrived. The prediction was already logged. That film was off by 0.280. Average miss moved from 0.286 to 0.285. Rank agreement moved from 0.762 to 0.750.

The accuracy record is the company. Integrity is the product.

Methodology questions: iscreeningroom.com