Post

Reconstructing AI Agent Activity Two New Scripts for Forensic Review

Reconstructing AI Agent Activity Two New Scripts for Forensic Review

Reconstructing AI Agent Activity: Two New Scripts for Forensic Review

A major update to FOR577 has added new material on investigating AI usage in incident response, discussing 8 of the most popular AI coding assistants and agents including Claude Code, Codex, Gemini CLI, Cursor, Copilot, Warp, Windsurf, and Qwen Code. To aid in forensic review, two new Python scripts have been added to a GitHub repository for examining the records left behind by OpenCode and Hermes: opencode-chat-replay.py and hermes_forensic_extract.py. The motivation behind these tools was to locate chat history and other evidence for forensic analysis. 🚀

Both scripts focus on reconstruction, but they approach it differently. The opencode script turns stored session data into readable chat transcripts or structured exports, while the Hermes script extracts a broader collection of evidence, including conversation records, model usage, API request dumps, and application logs. These are forensic review tools, not tools for running or replaying an agent’s actions. The goal is to make recorded activity accessible for investigation, detailing what the user asked, what the assistant returned, what tool activity was captured, and what surrounding context remains available. Both scripts require Python 3.10 or later and use only the Python standard library. They allow use of --start and --end (in YYYY-mm-dd [HH[:MM[:SS]]] format) to narrow the time range of the extraction, and -f/--file or -d/--dir to point to where the evidence can be found, e.g. from a mounted image or a tarball collected from a system under investigation.

Specifically, opencode-chat-replay.py reconstructs session transcripts from opencode’s SQLite database, located by default at: ~/.local/share/opencode/opencode.db. The script supports both the separate message and part tables and the newer consolidated session_message storage. This distinction matters because the transcript is more than a list of text messages; stored turns can contain text, assistant reasoning, tool calls, completion information, errors, and token or cost accounting. When listing sessions, the output includes session IDs, slugs, message and part counts, creation times in UTC, and titles. Markdown is the default output format for this script, with each transcript starting with session metadata, including the session ID, slug, working directory, version, timestamps, cost, tokens, and turn count. The conversation follows in numbered role sections, and reasoning and tool calls appear in collapsible < details> blocks, keeping the main transcript readable while retaining access to the supporting material. For structured analysis, JSON output contains a session object and an array of turns, while JSONL output places one turn on each line, with session metadata embedded in each record.

Critically for forensic integrity, before querying SQLite, both scripts copy the SQLite database and any present -wal and -shm sidecars into a temporary directory. This works around a potential issue in opening the db read-only, that still might result in a write to the database due to partially staged results when we don’t want to potentially tamper with the evidence. The scripts open that snapshot with SQLite’s read-only URI mode and remove the temporary directory on exit. This ensures that the process does not write to or delete files in the original opencode data location.

Read full article

This post is licensed under CC BY 4.0 by the author.