06 - Oct 2026

GenAI-assisted Digital Forensics Weaknesses and Mitigations in SOLVE-IT

“To do more faster, cyber investigators are adopting AI and large language models (LLMs), not realizing that these systems lack requisite knowledge. Without the expertise in planning an investigation, circumventing known problems, evaluating digital evidence, and conveying forensic results, increased speed raises the risk of unseen errors. Applied carelessly, AI is an amplifier, accelerating our existing flaws while adding new ones. Applied responsibly, AI can solve persistent challenges and ensure epistemic security in cyber investigations.” (Eoghan Casey, Ensuring Epistemic Security in AI-Driven Cyber Investigations, Blog@CACM)

Since 2023, the DFRWS community has been developing a systematic approach to filling knowledge gaps and enhancing quality assurance in digital forensics. This work is codified in the open source SOLVE-IT knowledge base and supporting tooling, cataloguing potential weaknesses in digital forensics techniques and associated mitigations. SOLVE-IT is a DFRWS supported project that covers the full digital forensics lifecycle and formalizes Error Mitigation Analysis to establish confidence in digital and multimedia evidence (ASTM E3016-18). For instance, the SOLVE-IT generate_evaluation.py Python program generates a worksheet for selected combinations of techniques to help practitioners systematically record how the potential weaknesses in their digital forensics workflow have been addressed.

The need for such systematic solutions increases as GenAI is used more to query case data at speed and scale. The independent investigation into the Hugging Face incident highlights ramifications of heavily delegating analysis to “often-unreliable AI agents.” The community recently updated SOLVE-IT with weaknesses of GenAI-assisted tools along with associated mitigations, as detailed in the technical article “Adding genAI-assisted techniques to SOLVE-IT.” The article defines the technique of using an AI-based prompt for interrogating digital evidence, and demonstrates how we systematically enumerate associated weaknesses and mitigations using the TRWM workflow.

Weaknesses include incomplete, inaccurate, and misinterpreted results when using GenAI to find artefacts, reconstruct events, discover links, or summarize data. Causes of errors in GenAI results are addressed, including missing pertinent information because the context window was exceeded or the search stopped when something else relevant was found. GenAI-assisted tools can make an incorrect inference, claiming that an event occurred that did not, or incorrectly asserting that something happened at a specific time or place. When a GenAI-assisted capability produces wrong results, this can have grave consequences in a digital forensics investigation, potentially implicating an innocent individual or excluding an offender from further inquiries.

Mitigating certain weaknesses in GenAI-assisted tools can be challenging, and are a shared responsibility between the tool developers and users. Robust guardrails are needed to control access to data and systems, avoid presenting inferences as facts, support transparency and reproducibility, and improve accuracy and reliability. Significant work is needed in this area to guide the design of safe and reliable GenAI-assisted digital forensics tools.

If you have weaknesses for this technique and mitigations that can be demonstrated to be effective, please submit to SOLVE-IT via GitHub directly or use the ‘suggest an edit’ button within the SOLVE-IT Explorer.