A scam caller only needs three seconds of your mother's voice, maybe lifted from a video she posted online, to clone it well enough to fool you. That's not a hypothetical, it's how a lot of 2026's phone scams actually work now. But voice cloning is just one of four different tricks hiding under the single word "deepfake," and each one is made differently, leaves different clues, and needs a different kind of check.
The four tricks behind every deepfake
Synthetic voice, or voice cloning, only needs a short sample of someone's real voice, sometimes just a few seconds, pulled from a video, a voicemail, or a social media post. An AI model picks up the tone, pace, and accent, then generates brand new speech in that same voice, saying whatever the scammer types. This is the trick behind most impersonation scam calls right now, since a cloned voice is finally good enough to fool a family member who isn't expecting it.
Face-swapping takes real video and replaces one person's face with someone else's, frame by frame, while keeping the original body, background, lighting, and movement the same. It's the oldest of the four tricks. And because it's built on top of real footage, it's often the easiest one to investigate, the original clip might still be sitting online somewhere if you look.
Image generation works differently. Instead of changing something real, it builds a photo from nothing but a text description. There's no original photo hiding underneath it, so there's nothing to trace back to. That's exactly what makes fake "photos" of events that never happened so hard to debunk with a reverse image search alone.
Full video manipulation is the newest and hardest one to catch. Instead of swapping a face onto footage that already exists, these tools generate an entire clip from scratch, the motion, the lighting, the audio, all of it, with nothing real underneath any part of it. Once tools like Sora 2 and Veo 3.1 made this mainstream, the question stopped being "was this face swapped?" and became "did this footage exist at all?"
How to actually check what you're looking at
Each trick has a check that fits it best. If you're unsure about a voice, don't try to listen harder, just hang up and call the person back on a number you already have saved. Or ask something a script couldn't have prepared for, since a clone can't improvise around a question it wasn't fed in advance. For a face-swapped video, a reverse image or video search actually helps, since the original footage the fake was built from might still be sitting online, unaltered. For a generated image, look at the small details a model still tends to get wrong: hands, background text, logos, reflections, and shadows that don't quite match up. For a fully generated video, mute it and just watch the face and body language, since the audio and lip movement are usually generated separately and can drift out of sync in ways that are much easier to catch once you can't hear it.
None of these checks are foolproof anymore, though. A proper detection tool, the kind that gives you a probability across several signals instead of a flat yes-or-no, is a genuinely useful second opinion for anything that actually matters. Pair that with checking whether the file carries a C2PA or SynthID credential, and you've got a decent process.
Why detection keeps getting harder
This isn't about people getting worse at spotting fakes. It's that the tools making them got dramatically better, and fast. Just a few years ago, careless mistakes, warped hands, mismatched blinking, a slightly robotic voice, gave a fake away to anyone paying attention. Now, studies show people can only correctly spot a good AI-generated video or image about a quarter of the time. And voice cloning has crossed what researchers call the indistinguishable threshold, meaning even trained listeners can't reliably tell a clone from the real thing in normal conditions anymore.
Detection tools run into the same problem from the other side. NIST's research on deepfake forensics has found that methods which work well in a controlled lab test often do noticeably worse on real content, where compression, re-uploading, and whatever processing a platform applies have already wiped out the small clues a detector would otherwise catch. It's a moving target on both sides. As the generators keep improving, the signals that used to give a fake away keep shrinking, which is exactly why relying on a sharp eye alone stopped being a reasonable plan.