Authors: Michael Rettinger, Ben Beaumont, Nhien-An Le-Khac, Hong-Hanh Nguyen-Le
DFRWS APAC 2026
Abstract
Deepfake, AI-generated, and conventionally tampered images increasingly appear in fraud, harassment, misin formation, and other investigations in which investigators must assess whether a questioned image is authentic, manipulated, or synthetic. Although deepfake-detection research is extensive, much less is known about the forensic reliability of the public tools that investigators, journalists, and first responders can access during early-stage triage. This paper evaluates six publicly accessible tools from two operational paradigms: analyst-interpreted forensic platforms (InVID & WeVerify, FotoForensics, and Forensically) and automated AI classifiers (DecopyAI, FaceOnLive, and Bitmind). Using blinded protocols, Dataset 1 assessed the forensic platforms on 100 authentic, tampered, and AI-generated images, while Dataset 2 assessed the classifiers and one blinded investigator on 150 real and deepfake images spanning facial and scene categories. Within their respective datasets, the two paradigms showed distinct error patterns. Forensic-platform assessments achieved high sensitivity but sometimes interpreted traces from genuine or benignly processed images as evidence of suspicious manipulation, whereas AI classifiers maintained high specificity for real images but missed substantial numbers of deepfakes, including all sampled HeyGen outputs. On Dataset 2, the investigator achieved a higher accuracy point estimate than every classifier. Human–AI disagreement was strongly asymmetric: most disagreements occurred when the investigator correctly identified a deepfake that a classifier labelled real. These findings indicate that publicly accessible image-authentication and deepfake-detection tools are useful for early-stage triage but should not be treated as standalone evidence. We recommend multi-tool usage, documented human review, and escalation of disagreement cases.