AI/ML Security Research

Kian Esmaeili builds systems that catch what security tools miss.

His research asks whether the metrics defensive AI reports about itself can be trusted — and builds the endpoint telemetry, evaluation methods, and embedded-security tooling to find out.

First-author, IEEE SoutheastCon 2026 · Incoming M.S. Cybersecurity, Georgia Institute of Technology


Mission

Most detection systems trust their own metrics. That trust is usually unearned.

Endpoint Behavior

Sysmon telemetry, evaluated the way an attacker experiences it — over time, not one event at a time.

Embedded Exploitation

Side-channel and fault-injection attacks against firmware, where security has a physical signal, not just a log line.

Applied AI Systems

Models built to be checked, not just trusted — where the evaluation is as rigorous as the architecture.


Research Journey

How the thinking evolved

Federal Reserve Bank of Atlanta
Systems change decisions before people do

Building automation for cybersecurity examinations showed how much of enterprise risk assessment depends on data most people never see directly. It set the question that still drives this work: how do you build a system that improves a decision before a human is asked to make it?

IEEE SoutheastCon 2026
Metrics can lie

Studying the 2024 CrowdStrike Falcon outage led to a first-author paper on AI-driven update validation — and to a harder problem underneath it: most ML-security results report accuracy numbers that don't survive contact with real-world behavior.

Temporal Collapse
Behavior, not events

Benchmarking ransomware detection on 6,200+ labeled Sysmon events, event-level models looked almost perfect while behavioral aggregation told the truth — 75–82%, not 99%. That gap now has a name, "temporal collapse," and a paper under review at EAI Endorsed Transactions.

Georgia Tech Research Institute
Security has a physical layer

Side-channel and fault-injection work on embedded firmware, using the ChipWhisperer-Nano platform, showed that the same discipline — question the signal, not the summary — holds below the operating system, not just above it.

Operation North Guard
A paper should be explorable

Static PDFs hide the most interesting part of research: how a result was reached. A capture-the-flag platform built to teach that reasoning became the basis of the next paper.

Living Papers
Read the paper by running it

That idea became a real project: turning published papers into interactive investigations, where the reader tests the hypothesis and watches the evidence collapse or hold. The first Living Paper — built on the Sysmon-ML temporal-collapse result — is live.

Explore Living Papers →
What's Next
Defense that can explain itself

The direction: systems that combine behavioral telemetry with AI reasoning to detect threats earlier — and hold up when someone asks why the system made that call.


Publications

Publications

The Sysmon-ML paper below is now also a Living Paper — an interactive investigation instead of a static PDF, built on its own findings.

AI-Driven Update Validation in Endpoint Security

Why AI-driven validation could have caught the failure mode behind the 2024 CrowdStrike Falcon outage before it shipped.

Read on IEEE Xplore →
Published IEEE SoutheastCon 2026 First Author · IEEE Xplore
Operation North Guard: A Multi-Variant Dynamic Flag Framework for Cybersecurity Education

A per-student, salted-hash flag architecture that makes cybersecurity coursework resistant to answer-sharing without adding grading overhead.

Under Review ISCAP 2026 First Author
Misleading Performance in Sysmon-Based Machine Learning

Why near-perfect ransomware-detection accuracy is usually a measurement artifact of event-level representation, not a real result — and how temporal structure exposes it.

Explore the Living Paper →
Under Review EAI Endorsed Transactions First Author

Case Studies

Engineering case studies

Not a list of technologies — the problem, the constraint, and the decisions that followed.

Operation North Guard — Cybersecurity CTF Platform

Next.js · React · TypeScript
Problem
Academic capture-the-flag exercises get their flags shared and re-used across semesters, which quietly breaks their value as an assessment.
Approach
A multi-variant dynamic flag system, salted with SHA-256 per student, layered on a modular MDX challenge engine spanning network forensics, cryptography, reverse engineering, SQL injection, and cloud security.
Outcome
10+ challenges shipped; the flag architecture became the basis of a first-author ISCAP 2026 paper, now under review.
Live platform → Source →

Living Papers — Interactive Research Platform

Interactive Publishing
Problem
Published research is skimmed, not explored — its strongest evidence locked inside static figures and dense prose.
Approach
A format — hook, explore, validate — that turns a paper's findings into hypotheses the reader tests directly, with every interaction anchored to the exact published figure.
Outcome
The first Living Paper is live, built on the Sysmon-ML temporal-collapse result: flip the switch and watch a 100%-accurate detector collapse to its honest number.
Visit Living Papers →

AI-Powered Security Assistant

n8n · LLM · RAG
Problem
Research notes and reference material accumulate faster than they can be searched by hand.
Approach
An n8n-orchestrated pipeline pairing a vector database (RAG) with an LLM, exposed through Telegram for low-friction, real-time interaction.
Outcome
A working ingestion → semantic search → generation pipeline for personalized knowledge retrieval.

Credentials

Credentials

Education

Georgia Institute of Technology

May 2028

M.S. Cybersecurity (Information Technologies), Atlanta, GA · starting August 2026

University of North Georgia

May 2026

B.S. Cybersecurity, Minor in Entrepreneurship · GPA 3.82/4.00 · Honors Thesis on AI-driven endpoint security

Recognition
  • 2x recognized, NSA/DHS Codebreaker Challenge
  • 2x Winner, UNG Pitch Challenge
  • National Cyber Scholarship (SANS / CyberStart America) — national high scorer
  • University Innovation Fellow — Stanford d.school; presented at the international UIF conference, University of Twente
  • 3rd of ~38 teams, Federal Reserve System Board AI Tournament
  • UNG Representative, Virginia Tech & VMI Cyber Immersion Camp

Technical Foundation

Technical foundation

Programming & Systems
PythonSQLPowerShellBashJavaScriptLinuxKali LinuxWindows
Detection & Forensics
SysmonEndpoint TelemetryMalware AnalysisNetwork SecurityWiresharkNmapFTK Imager
ML & Reverse Engineering
Model EvaluationBehavioral ModelingGhidraIDA Prox64dbgBinary Ninja

Contact

How we could work together

Open to research collaboration, cybersecurity and AI/ML security roles, and conversations about problems worth solving.