Analyzing packet captures manually takes time. Lots of it.
LLMs can help, but they’re not magic.
What they’re actually good at
Pattern recognition - Spotting HTTP requests that look suspicious, identifying unusual timing patterns, flagging protocol violations.
Correlation - Connecting related packets across a capture, finding conversations buried in noise.
Initial triage - Rough categorization of traffic types so you know where to focus.
What they’re bad at
Encryption - Obviously. LLMs can’t decrypt traffic.
Binary protocols - They struggle with raw binary data. Pre-process it first.
False positives - They flag lots of suspicious-looking but legitimate traffic. You still need to verify.
Practical example
You have a 500MB pcap file. Thousands of connections. Something in there is malicious.
Manual approach: Hours of filtering, following TCP streams, checking DNS queries.
LLM-assisted approach:
- Extract metadata (IPs, ports, timing, packet sizes)
- Feed to LLM: “Find connections with unusual timing patterns”
- LLM narrows it down to 50 suspicious connections
- Manual analysis of those 50
- Find the actual malicious traffic
Result: 15 minutes instead of hours.
Simple Python implementation
import pyshark
import anthropic
def analyze_capture(pcap_file):
# Parse pcap
cap = pyshark.FileCapture(pcap_file)
# Extract connection metadata
connections = {}
for packet in cap:
if 'IP' in packet:
src = packet.ip.src
dst = packet.ip.dst
key = f"{src}-{dst}"
if key not in connections:
connections[key] = {
'packets': 0,
'bytes': 0,
'timestamps': []
}
connections[key]['packets'] += 1
connections[key]['bytes'] += int(packet.length)
connections[key]['timestamps'].append(float(packet.sniff_timestamp))
# Build summary
summary = []
for conn, data in connections.items():
intervals = []
times = data['timestamps']
for i in range(1, len(times)):
intervals.append(times[i] - times[i-1])
avg_interval = sum(intervals) / len(intervals) if intervals else 0
summary.append({
'connection': conn,
'packets': data['packets'],
'bytes': data['bytes'],
'avg_interval': avg_interval
})
# Send to Claude
client = anthropic.Anthropic()
prompt = f"Analyze these network connections. Flag anything suspicious:\n\n{summary}"
response = client.messages.create(
model="claude-3-5-sonnet-20241022",
max_tokens=2000,
messages=[{"role": "user", "content": prompt}]
)
return response.content[0].text
# Run it
results = analyze_capture("capture.pcap")
print(results)
This is basic. It works.
Beaconing detection
Malware often beacons home at regular intervals. LLMs can spot this:
Connection A: packets every 60s ± 2s
Connection B: packets every 14m ± 30s
Connection C: random timing, 200-4000s gaps
Connections A and B look like beacons. Connection C looks normal.
The LLM can calculate interval consistency and flag the regular ones.
DNS tunneling
Suspicious DNS patterns:
- Lots of TXT record queries
- Subdomains that look like base64
- High query frequency to one domain
- Response sizes way larger than typical
LLM prompt: “Analyze these DNS queries. Find patterns suggesting tunneling.”
It’ll flag the weird ones. You verify if it’s actual tunneling or just a poorly designed app.
Limitations
LLMs hallucinate. Sometimes they’ll confidently tell you traffic is malicious when it’s not.
Always verify their findings against the actual packets.
They’re pattern matchers, not security analysts. Use them to narrow your search space, not make final decisions.
Data privacy
You’re potentially feeding sensitive network data to an API.
Options:
- Strip sensitive info before sending (IPs, usernames, actual payload content)
- Use local models (less capable but private)
- Only analyze metadata, never full packet contents
If it’s truly sensitive traffic, don’t send it to external APIs at all.
Cost
API calls add up. A 500MB pcap might generate $2-5 in API costs depending on how much you send.
Batch your analysis. Send summaries not raw data. Cache results.
Resources
Download toolkit - Scripts for packet analysis with LLM assistance