Hello, I'm

Amirreza Vishteh

AI Safety & LLM Security Researcher

I study how large language models can be secretly backdoored — and build the tools that find, audit, and explain those hidden behaviours.

  • M.Sc. Computer Engineering · Sharif University of Technology
  • DML Lab · AI safety research
  • Tehran, Iran
Portrait of Amirreza Vishteh
NowLLM backdoor detection
Top 0.3%National entrance exam
National entrance exam (#475)
Top 0.3%
Research write-ups
15+
Open-source repos
20+
Courses as TA
6

About me

Making AI systems safe to trust

I'm an AI safety researcher focused on trustworthy AI and large language models. I'm pursuing my M.Sc. in Computer Engineering at Sharif University of Technology, where I'm a member of the Data Science and Machine Learning Laboratory.

Most of my research is about backdoors in LLMs: planting them in controlled settings, stress-testing state-of-the-art scanners against harder variants, and building detectors that work without knowing the trigger in advance — from attention-based signals to small watchdog models for LoRA adapters. I care as much about whether a detector explains the backdoor correctly as about whether it raises a flag.

Before Sharif, I earned my B.Sc. in Computer Engineering at Iran University of Science and Technology (IUST), where I worked on Persian NLP and co-authored a mental-health chatbot paper published at AbjadNLP 2025. I've also served as a teaching assistant for courses from Stochastic Processes and Machine Learning to NLP and Computer Architecture, and I build applied healthcare-AI tools such as DentalMind.

Research interests

AI SafetyTrustworthy AILLM Backdoor DetectionLarge Language ModelsInterpretabilityHealthcare AIPersian NLPSpeech Processing

Research

Selected publications

All publications

Open source

Featured projects

All projects
AI Safety & SecurityNov 2025

Online Backdoor Detector for LLMs

A “neuro-statistical” watchdog that flags backdoored LoRA adapters without knowing the trigger, using the suspect model’s confidence and semantic-consistency signals and a multi-head verdict: clean, suspicious, or backdoor detected.

LLM backdoorsLoRAdetection+1
Python Write-up
AI Safety & SecurityAug 2025

BAIT Weakness Zoo

Extends the BAIT (S&P 2025) LLM backdoor scanner with sparsemax candidate selection and a controlled zoo of benign, standard, and evasive backdoored LoRA models to locate exactly where the scanner fails while the backdoor still fires.

backdoor scanningBAITLoRA+1
Python Write-up
AI Safety & SecurityNov 2025

AttentionGuard

A backdoor detection system for LLMs based on attention analysis, introducing a “self-centeredness” score to catch sophisticated backdoors that evade confidence-only detectors.

attention analysisLLM backdoorsdetection
Python
Computer Vision & Healthcare AI2024 – Present

DentalMind

Trustworthy AI for dental diagnostics: per-tooth radiograph analysis designed to support, not replace, a clinician’s judgment. Python, PyTorch, computer vision, FastAPI and React.

healthcare AIcomputer visionFastAPI+1
AI Safety & SecurityAug 2026

Devign Reproduction & Leakage Audit

An end-to-end PyTorch reproduction of Devign (NeurIPS 2019) on the authors’ released data — composite AST/CFG/DFG graphs, gated graph layers and the Conv readout — plus a commit-disjoint leakage audit of the benchmark.

graph neural networksvulnerability detectionreproducibility
Python Write-up
Networks & SystemsFeb 2026

MECP-GAP

Implementation of mobility-aware MEC planning with GNN-based graph partitioning (IEEE TNSM 2024): assigning 5G base stations to edge servers to minimise handover cost under load-balancing constraints.

graph neural networks5Gedge computing+1
Python Write-up

Gallery

Moments

Lab life, campus, friends, mountains — and two very good dogs.

Open gallery

Contact

Let's work together

I'm always happy to talk about research collaborations in AI safety and trustworthy AI, open-source work, and interesting machine-learning problems. The fastest way to reach me is by email.