Skip to project
← Projects

PROJECT_05 / Applied security research

AI4Security Research

Testing where AI helps security analysis—and where it falls short.

STACKLLMs · Azure · PCAP analysis · Modbus · PostgreSQL

01

Overview

The research had two tracks: enriching security alerts with asset and CVE context, and analyzing industrial PCAPs from the 2023 UNB CIC Modbus dataset. The question was whether AI could support useful security analysis rather than simply generate plausible explanations.

My contribution

I worked with a co-op student on model evaluation, packet-data preparation, alert enrichment, and the design of a validation loop for AI-generated analysis.

02

The problem

Raw packet bytes require exact interpretation. Models struggled with byte counting and hex-level reasoning, while narrow training examples did not generalize reliably to varied traffic. Convincing output was not sufficient evidence of correct analysis.

03

Engineering approach

We evaluated open-source cybersecurity-tuned models, prepared structured JSON for Azure fine-tuning, and compared prompting approaches. Traffic was reformatted into higher-level protocol and NetFlow-style context to give models a clearer basis for reasoning.

  1. 01CIC Modbus PCAP
  2. 02Protocol / flow context
  3. 03Model comparison
  4. 04Output validation

Fine-Tuning on Structured PCAP Data

Fine-tuned models performed efficiently when the incoming PCAP closely matched the traffic structure seen in the training examples, but the approach degraded when the packet patterns varied from the baseline.

Prompt-Based Reasoning

One-shot, multi-shot, and chain-of-thought prompting handled variation better than fine-tuning alone, but the outputs were still not accurate enough to be trusted as a reliable threat-detection workflow.

Agentic Baseline + Validation

A more agentic pipeline showed better potential by having models generate expected frames or baselines, compare them with observed traffic, and use a feedback loop to reject and regenerate invalid analyses.

04

Alert enrichment

Dummy asset records in PostgreSQL were combined with CVE information so an LLM could explain potential vulnerabilities in the context of an asset. This explored an assistive workflow rather than an autonomous detection decision.

05

Technical challenges

01

Exact packet interpretation

Problem
Models were unreliable at byte counting and raw hex interpretation.
Approach
Preprocess traffic into structured protocol and flow context.
Result
Shift the task toward reasoning over meaningful evidence.
02

Handling variation

Problem
Fine-tuning worked best on narrow, repeatable traffic structures.
Approach
Compare prompt-based reasoning with fine-tuned models.
Result
Observe greater flexibility, but insufficient detection-grade accuracy.
03

Controlling invalid output

Problem
A single model response could be plausible but wrong.
Approach
Design a baseline-generation and validation loop with rejection and correction.
Result
Identify a more controlled direction for future research.
06

Result

The models were not reliable enough to serve as a standalone threat-detection engine. Structured inputs and validation loops were more promising than direct raw-byte reasoning. The project was put on hold after the co-op term ended.

07

Lessons learned

  • Input representation strongly influences the usefulness of AI analysis.
  • Flexible reasoning and reliable detection are different requirements.
  • Validation must be part of the workflow, not an assumption about the model.