Packet Analysis with LLMs

Analyzing packet captures manually takes time. Lots of it.

LLMs can help, but they’re not magic.

What they’re actually good at

Pattern recognition - Spotting HTTP requests that look suspicious, identifying unusual timing patterns, flagging protocol violations.

Correlation - Connecting related packets across a capture, finding conversations buried in noise.

Initial triage - Rough categorization of traffic types so you know where to focus.

What they’re bad at

Encryption - Obviously. LLMs can’t decrypt traffic.

Binary protocols - They struggle with raw binary data. Pre-process it first.

False positives - They flag lots of suspicious-looking but legitimate traffic. You still need to verify.

Practical example

You have a 500MB pcap file. Thousands of connections. Something in there is malicious.

Manual approach: Hours of filtering, following TCP streams, checking DNS queries.

LLM-assisted approach:

  1. Extract metadata (IPs, ports, timing, packet sizes)
  2. Feed to LLM: “Find connections with unusual timing patterns”
  3. LLM narrows it down to 50 suspicious connections
  4. Manual analysis of those 50
  5. Find the actual malicious traffic

Result: 15 minutes instead of hours.

Simple Python implementation

import pyshark
import anthropic

def analyze_capture(pcap_file):
    # Parse pcap
    cap = pyshark.FileCapture(pcap_file)
    
    # Extract connection metadata
    connections = {}
    for packet in cap:
        if 'IP' in packet:
            src = packet.ip.src
            dst = packet.ip.dst
            key = f"{src}-{dst}"
            
            if key not in connections:
                connections[key] = {
                    'packets': 0,
                    'bytes': 0,
                    'timestamps': []
                }
            
            connections[key]['packets'] += 1
            connections[key]['bytes'] += int(packet.length)
            connections[key]['timestamps'].append(float(packet.sniff_timestamp))
    
    # Build summary
    summary = []
    for conn, data in connections.items():
        intervals = []
        times = data['timestamps']
        for i in range(1, len(times)):
            intervals.append(times[i] - times[i-1])
        
        avg_interval = sum(intervals) / len(intervals) if intervals else 0
        
        summary.append({
            'connection': conn,
            'packets': data['packets'],
            'bytes': data['bytes'],
            'avg_interval': avg_interval
        })
    
    # Send to Claude
    client = anthropic.Anthropic()
    prompt = f"Analyze these network connections. Flag anything suspicious:\n\n{summary}"
    
    response = client.messages.create(
        model="claude-3-5-sonnet-20241022",
        max_tokens=2000,
        messages=[{"role": "user", "content": prompt}]
    )
    
    return response.content[0].text

# Run it
results = analyze_capture("capture.pcap")
print(results)

This is basic. It works.

Beaconing detection

Malware often beacons home at regular intervals. LLMs can spot this:

Connection A: packets every 60s ± 2s
Connection B: packets every 14m ± 30s
Connection C: random timing, 200-4000s gaps

Connections A and B look like beacons. Connection C looks normal.

The LLM can calculate interval consistency and flag the regular ones.

DNS tunneling

Suspicious DNS patterns:

  • Lots of TXT record queries
  • Subdomains that look like base64
  • High query frequency to one domain
  • Response sizes way larger than typical

LLM prompt: “Analyze these DNS queries. Find patterns suggesting tunneling.”

It’ll flag the weird ones. You verify if it’s actual tunneling or just a poorly designed app.

Limitations

LLMs hallucinate. Sometimes they’ll confidently tell you traffic is malicious when it’s not.

Always verify their findings against the actual packets.

They’re pattern matchers, not security analysts. Use them to narrow your search space, not make final decisions.

Data privacy

You’re potentially feeding sensitive network data to an API.

Options:

  • Strip sensitive info before sending (IPs, usernames, actual payload content)
  • Use local models (less capable but private)
  • Only analyze metadata, never full packet contents

If it’s truly sensitive traffic, don’t send it to external APIs at all.

Cost

API calls add up. A 500MB pcap might generate $2-5 in API costs depending on how much you send.

Batch your analysis. Send summaries not raw data. Cache results.

Resources

Download toolkit - Scripts for packet analysis with LLM assistance