Figures as of August 2026. Regenerated as the corpus grows. Written for filmmakers and skeptical readers alike - plain meaning first, every number kept, limitations stated alongside results.
Living document · updated with the corpus
aiScreeningRoom predicts how real-world audiences will receive an unreleased film by watching the actual footage - the entire film - with no human test audience.
That claim rests on two independently checked links. Neither ever peeks at the answer it is graded on:
No step in the chain has access to the answer it is being graded on.
Panels are nationwide audiences selected for interest in the film’s genre, premise, and comparable titles - a taste-matched sample of the film’s natural audience, not whoever showed up to a theater on one night.
Panel testing is ongoing; the corpus and every figure in this document grow with it.
For every corpus film that went on to earn an honest organic public IMDb rating - at least 100 votes, identity verified by exact ID (including retitled releases), a ballot histogram free of organized voting, and a panel of at least 100 respondents - we compared the panel score against the rating the public eventually gave. Each film is graded by a mapping that never saw it.
Across the 36 qualifying films:
27 of 36 (75%) within three-quarters of a point; 23 of 36 (64%) landed within half a point. Under a stricter bombing screen, that ranking agreement rises to ρ = 0.70. No film is ever graded by a mapping that saw it, and the mapping preserves the scale’s spread - it can call a film exceptional or a disaster rather than compressing every prediction toward the middle.
A single IMDb mean can hide review bombing and fan stuffing. We record the share of ballots at 1 star and at 10 stars and exclude films where either pole holds 30% or more of all votes. Worked example: one politically charged faith film shows IMDb 1.6 from a histogram that is 71% one-star votes, against a panel score of 4.42 - that is protest voting, not audience reception. Without the screen, correlation drops to 0.40 - and the misses are dominated by films with organized-voting histograms, not by panel error. Public ratings are a noisy, manipulable proxy; a several-hundred-person taste-matched panel is a cleaner read of the same construct. The screen keys on a single pole. Some pages that pass it still show a U-shape - elevated shares at both one star and ten stars - so the published mean is pulled by the poles while the middle of the ballot (2-9) is closer to ordinary taste. On the same 36 films, using that middle-of-ballot mean only for titles meeting a pre-stated U-shape rule, 34 of 36 (94%) land within one point.
IMDb’s own fraud detection visibly discounts distrusted ballots - the spread between a film’s raw ballot mean and its published rating. Films carrying large discounts have their influence on the mapping scaled down in proportion to that evidence (no film is deleted by it). The screen keys only on IMDb-internal evidence, never on disagreement with the panel.
On 9 sealed holdout films - locked away from training and tuning; never trained on, never tuned on - graded against real panel results (3 viewing replicates per film, average reported):
Slightly conservative on average (bias −0.116) - the misses ran low, not random.
A skeptic will ask: is this just recognizing famous movies? We tested that. We asked, with no footage, what the system already knew about a film’s public reception. On 44 films it could not recall, it was about as accurate as on films it could recall (off by 0.26 vs 0.28 on the 1 to 5 scale). So the score is coming from the picture, not from the title.
Those 44 still include films used while we built the method. The accuracy number is the 9 films locked away from that work. That is why 0.285 on 9 films is the claim, not 0.26 on 44.
Unreleased films don’t have public data - which is the entire reason a prediction is needed. The prospect-facing claim therefore composes two independently measured links: AI→panel (sealed holdout) and panel→public (bridge study above). The AI is never calibrated on public ratings - and each bridge film is graded by a mapping that never saw it.
Production runs the v2.3t viewing-pass prompt (band-commitment scoring). The 9-film figures above (average miss 0.285, ρ 0.750) were produced under that v2.3t prompt. Buyer-facing accuracy claims rest on Link 1 and Link 2 above - the sealed holdout for AI→panel, not a re-score of films used to build the method. The licensed holdout look for this operating prompt is spent; further method work requires a newly licensed look.
Every beta film is eligible for a prospective accuracy ledger: prediction logged while no panel actual exists, frozen at log time, scored only when panel results later arrive - never rewritten. The sealed 9-film figures above stay as published. When a beta film later has panel results, that row is scored on the prospective ledger. We will report that larger set as a separate living figure. It does not rewrite the sealed 9. Existing corpus evaluations are refused as non-prospective. Your title is never published. No film, report, score, or excerpt is identified publicly without your prior written consent. Published figures are aggregate.
When a published figure on this page changes, the prior number stays in the log below.
The accuracy record is the company. Integrity is the product.
Methodology questions: iscreeningroom.com