AI Video Detectors: Fighting Deepfakes in the Sora Era
For most of the internet's history, video carried a quiet authority that text and even photographs never had. A written claim could be fabricated in seconds; a photo could be Photoshopped by anyone with patience. But moving footage of a person saying and doing something felt like the closest thing to being in the room. That assumption is now dead. Between convincing face-swap deepfakes and text-to-video models like OpenAI's Sora, Runway's Gen models, Google's Veo, and the wave of successors chasing them, we have crossed into an era where a photorealistic clip of an event that never happened can be produced by someone with no skill, no budget, and no footage to start from. AI video detectors are the tools built to push back against that shift, and this guide is an honest account of what they can do, what they cannot, and why the people building them will tell you privately that they are losing.
The stakes are not abstract. Synthetic video is already being used to drain bank accounts, swing the emotional temperature of elections, humiliate private individuals through non-consensual sexual imagery, and manufacture "evidence" that never happened. A finance worker at a multinational was tricked into wiring millions after joining a video call where every other participant, including the CFO, was a deepfake. Fabricated clips of politicians conceding, confessing, or inciting violence spread across social platforms faster than any fact-check can follow. The overwhelming majority of deepfake material circulating online is non-consensual pornography targeting women, most of them not public figures at all. Video detection is not a niche technical curiosity. It is one of the front lines of a much larger fight over whether recorded reality can still be trusted, and it is worth understanding exactly how thin that line has become.
Two very different problems wearing the same name
The word "deepfake" gets used for two fundamentally different things, and the distinction matters enormously for detection because the two require almost entirely different techniques to catch. Lumping them together is the first mistake most people make.
The first problem is the face-swap deepfake: you start with a real video of a real person and graft someone else's face, or a manipulated version of the original face, onto the existing footage. The body, the lighting, the room, the camera motion are all authentic recordings. Only the face and sometimes the voice have been synthesized and stitched in. This is the classic deepfake that emerged in 2017 and still powers most of the malicious material online, because it is cheap, it targets a specific real person, and the synthetic region is small and localized. The forgery lives at the seams — the boundary where the fake face meets the real neck, the mismatch between a synthesized mouth and the authentic jaw around it.
The second problem is fully generated video: there is no source footage at all. A model like Sora takes a text prompt and hallucinates every pixel of every frame from nothing — the people, the physics, the shadows, the reflections, the entire scene. This is the step-change that shattered the previous détente in the detection arms race. With face-swaps, a detector can hunt for the boundary between real and fake. With a fully generated clip, there is no boundary, because nothing in the frame is real. The whole image is synthetic and internally consistent in a way that face-swaps never were. Detectors trained for years on the tell-tale seams of face-swapping were suddenly confronted with footage that had no seams to find, and many of them simply fell over.
This is why you should be suspicious of any tool that claims to "detect deepfakes" without saying which kind. A detector tuned to spot blending artifacts around a swapped face may be nearly useless on a Sora clip, and a model trained to recognize the statistical fingerprint of a diffusion generator may miss a clean face-swap built with older methods. The problem space fractured, and the detectors fractured with it.
What a detector actually looks for
Underneath the marketing, AI video detectors are pattern recognizers hunting for the small ways synthetic footage still fails to imitate the messy physics of the real world. The specific tells have shifted over the years as generators improved, but they cluster into a few recognizable families, and knowing them helps you understand both the promise and the fragility of the whole enterprise.
Temporal inconsistency
A single frame of a modern deepfake can be flawless. The trouble for generators is making a thousand consecutive frames all agree with each other. Real video has a physical continuity that is genuinely hard to fake: an earring stays the same shape from frame to frame, a strand of hair follows gravity, a background object does not subtly warp when a head passes in front of it. Many detectors ignore individual frames almost entirely and instead study how pixels evolve over time, looking for the flicker, the shimmer, the objects that pop in and out of coherence. Fully generated video in particular has struggled with object permanence — a coffee cup that changes handles, fingers that merge and separate, text on a sign that dissolves into gibberish when you look between frames. Temporal analysis is one of the more durable detection strategies precisely because consistency across time is one of the hardest things for a generator to maintain.
The face that forgets to be human
Human faces do a lot of subtle involuntary things, and early deepfakes were notoriously bad at them. The famous one was blinking: generative models trained mostly on photographs, in which people's eyes are open, produced faces that blinked too rarely or at unnatural intervals. Detectors exploited this for a while, and then the generators learned to blink, and the tell evaporated. That cycle — a reliable artifact, a detector built on it, a generator that closes the gap, a dead detector — is the fundamental rhythm of this entire field, and it is worth burning into memory before you trust any single technique. Related tells include imperfect lip-sync, where the mouth shapes do not quite match the phonemes of the audio, unnatural gaze that does not track anything in the scene, and asymmetries in expression that the human visual system registers as "off" even when a viewer cannot articulate why.
Lighting and physics that do not add up
Generators reproduce the appearance of a scene without simulating the physics underneath it, so they routinely get the physics slightly wrong in ways a detector — or a careful human — can sometimes catch. Shadows fall in directions inconsistent with the visible light source. Reflections in eyes, glasses, or windows fail to match the surrounding environment. Skin catches light in a way that does not correspond to any coherent lighting setup. These same failures show up in still images, which is why video detection borrows heavily from the older discipline of image forensics; if you want the deeper version of that story, our companion piece on how AI image detectors work covers the pixel-level forensics that video inherits frame by frame.
Frame-level generator fingerprints
Every generation method leaves a statistical residue in the pixels — patterns in the frequency domain, subtle regularities in noise, artifacts of the upscaling and decoding steps a diffusion model uses to turn latent noise into an image. These fingerprints are invisible to the eye but detectable by a classifier trained to recognize the signature of a specific generator. The catch, and it is a serious one, is that these fingerprints are specific. A detector that has learned the fingerprint of one model may be blind to the next model, and blind to the one released next month, because each new architecture leaves a different residue. This is the deep structural weakness of fingerprint-based detection: it is always a description of the generators that already exist, never the ones about to arrive.
Biological signals hiding in the skin
The most striking detection approach exploits something generators have no reason to reproduce: the tiny color changes in a real person's skin caused by blood pulsing through the capillaries beneath it. Every heartbeat, your face flushes a fraction of a shade redder and then fades, far too subtly for the eye to see but measurable by a camera — a technique borrowed from remote photoplethysmography, the same principle behind contactless pulse monitors. Intel's FakeCatcher built its detector on exactly this idea: a genuine video of a person contains a coherent, physiologically plausible pulse signal across the face, and most synthetic video does not, because the generator was never modeling a circulatory system. It is a clever, almost poetic tell — the fake fails because it has no heartbeat. But it is not magic. It degrades on low-quality or heavily compressed footage where the subtle color signal is destroyed, and there is no guarantee future generators cannot be trained to fake a plausible pulse once it becomes a known target.
Your webcam is now an attack surface
For years, the reassuring assumption about deepfakes was that they took time and effort to produce — you could not fake a live, interactive human in real time, so a video call was a safe way to confirm someone was who they claimed to be. That assumption has collapsed. Real-time face-swapping software can now map a synthetic face onto a live camera feed with enough speed and quality to survive a video call, which means the person on the other end of your screen, moving and responding in real time, may not exist.
This turns everyday remote interaction into an attack surface. Job interviews conducted over video have been infiltrated by candidates using deepfake faces to disguise their real identity. Romance and investment scammers run live deepfake personas to build trust before asking for money. Most alarmingly, corporate fraud now includes fake video meetings where an employee is walked through an urgent wire transfer by what appears to be their own leadership team, every face on the call synthetic. Detecting a live deepfake is even harder than detecting a recorded one, because you cannot run a slow, careful forensic analysis on a stream you are watching in the moment. The practical defenses that actually work today are almost quaint: ask the person to turn their head fully to the side, where many real-time systems still break down; ask them to wave a hand in front of their face; or fall back on a code word or a second channel of confirmation. That the frontier of anti-fraud advice has come down to "ask the executive to show you their profile" tells you how thoroughly the ground has shifted.
Provenance: proving what is real instead of catching what is fake
Because detection is so fragile, a growing camp argues the whole strategy is backwards. Instead of trying to catch fakes after the fact — an endless, losing chase — we should cryptographically certify what is real at the moment of capture, and treat anything without that certification as unverified. This is the provenance approach, and it is philosophically the inverse of detection.
The leading standard is C2PA (the Coalition for Content Provenance and Authenticity), backed by camera makers, Adobe, Microsoft, and a broad industry coalition. The idea is that a camera or editing tool attaches a tamper-evident cryptographic manifest to a file — a signed record of where it came from and how it was edited — that travels with the content as "Content Credentials." A viewer can then check the credential and see a verifiable chain of custody. On the generative side, Google's SynthID embeds an imperceptible watermark directly into the pixels (and frames) of AI-generated content, including video from its models, so that a detector keyed to the watermark can flag the content as synthetic even after ordinary edits. Together these represent the "prove it's real" and "mark it as fake" halves of the provenance strategy.
Provenance is a genuinely better idea than detection in principle, because it does not depend on out-guessing an adversary who keeps improving. But it has hard limits. It only works when the capture device and the platform both support the standard, and most of the cameras and screenshots and re-uploads in the world do not. Metadata and watermarks can be stripped, either deliberately by a bad actor or incidentally by a platform that re-encodes everything on upload. A missing credential does not prove something is fake — it just means you do not know, which is true of almost everything online. And watermarks embedded by cooperative companies do nothing about the open-source generators that will happily produce unmarked video on request. Provenance raises the floor for honest actors and does very little to constrain determined dishonest ones. It is necessary, it is worth building, and it is not a solution on its own. The same tension between watermarking and detection plays out in audio, which our guide to AI music and voice detectors examines in detail — voice cloning is often the audio half of the same deepfake.
The platform problem
Almost no one runs a deepfake through a standalone detector. Synthetic video reaches people through platforms — social feeds, video sites, messaging apps — and so the real battle over AI video is being fought inside the content-moderation systems of a handful of enormous companies, at a scale that changes the nature of the problem entirely.
A tool like Deepware or a research detector can afford to spend seconds carefully analyzing a single clip. A platform ingesting hundreds of hours of video every minute cannot. Moderation at that scale has to be fast, cheap, and automatic, which means it leans on the least reliable end of detection and on provenance signals and watermarks where they exist. Platforms have responded mainly with labeling policies: YouTube and others now require creators to disclose realistic AI-generated or altered content and attach labels, and they honor provenance credentials where present. This shifts part of the burden onto disclosure and away from detection — a sensible move, since detection at platform scale is so unreliable, but one that depends on honesty from exactly the people most likely to be dishonest. The specific mechanics of how one major platform approaches AI content, both in video and in the channels that publish it, are worth understanding on their own, and we go deeper in our piece on YouTube channel AI detectors. The uncomfortable reality is that no platform can catch synthetic video reliably at the speed and volume it arrives, so a meaningful amount always gets through, and labeling policies are as much about managing that failure as preventing it.
Why compression quietly ruins everything
Here is the practical fact that undermines video detection more than any single generator improvement, and it is the one least discussed in the marketing. Detection depends on subtle pixel-level signals — the faint blood-flow color changes, the frequency-domain fingerprints, the small artifacts at a face-swap boundary. Every one of those signals is delicate. And every video that travels across the internet gets compressed, re-encoded, resized, and re-uploaded, often several times over.
When you upload a clip to a social platform, it does not store your file. It re-encodes it into its own format at its own bitrate, throwing away visual information the algorithm judges unimportant to a human viewer. Screen-recording a video and re-posting it compresses it again. A clip that has been shared, downloaded, re-uploaded, and re-shared across three platforms may have been through the compression grinder half a dozen times. Each pass smooths away exactly the faint statistical traces detectors rely on, while leaving the content perfectly convincing to a human eye. The result is a brutal asymmetry: compression destroys the evidence of forgery far faster than it destroys the forgery's persuasive power. A deepfake can survive a dozen re-encodings and still fool you completely, while the detector that could have caught it in the pristine original is left analyzing mush. This is why a tool that scores well on a clean research dataset can perform far worse on the messy, recompressed footage that actually circulates in the wild, and why laboratory accuracy numbers should be read with deep skepticism when the target is real-world video.
An honest verdict
If you take one thing from this guide, let it be this: video detection is currently, badly outpaced by video generation, and the gap is widening rather than closing. Every detection strategy in this article describes a fingerprint left by today's generators, and the generators improve continuously, closing each tell as fast as it is discovered. The economics are hopelessly lopsided — enormous, well-funded labs push generation forward, while detection is a comparatively small, reactive effort that is always responding to the last model rather than the next one. Compression erodes the signals detectors need, real-time deepfakes remove the time forensics require, and each new architecture leaves a different fingerprint that yesterday's detector has never seen. This is the same structural predicament that afflicts text detection, which we lay out in our foundational explainer on what an AI detector is and how detection works; the video version is simply harder, because footage is richer, heavier, and mangled more thoroughly on its way to your screen.
None of this means detectors are worthless. A good one can catch clumsy fakes, flag suspicious clips for a human to examine, and raise the cost and skill required to fool people — real value, even if it falls short of certainty. But you should never treat a detector's verdict on a video as proof, in either direction. A "real" score does not mean a clip is authentic; a "fake" score is not evidence a court or a newsroom should act on alone. The durable defenses are not technical at all: verify extraordinary footage through a second independent source, be suspicious of clips engineered to make you feel something urgent before you can think, and confirm identity over a channel that a deepfake cannot occupy. The honest position, the one this site holds across every medium, is that the tools are a fallible aid and never a final authority. Trust the process of verification, not the score — and when a video demands an immediate emotional reaction from you, treat that demand itself as the first thing worth doubting.