Psychological Health Chatbot, Detecting and Assisting Patients in their Path to Recovery
Proceedings of the 1st Workshop on NLP for Languages Using Arabic Script (AbjadNLP @ COLING 2025), pp. 64–77, ACL
Hello, I'm
AI Safety & LLM Security Researcher
I study how large language models can be secretly backdoored — and build the tools that find, audit, and explain those hidden behaviours.

About me
I'm an AI safety researcher focused on trustworthy AI and large language models. I'm pursuing my M.Sc. in Computer Engineering at Sharif University of Technology, where I'm a member of the Data Science and Machine Learning Laboratory.
Most of my research is about backdoors in LLMs: planting them in controlled settings, stress-testing state-of-the-art scanners against harder variants, and building detectors that work without knowing the trigger in advance — from attention-based signals to small watchdog models for LoRA adapters. I care as much about whether a detector explains the backdoor correctly as about whether it raises a flag.
Before Sharif, I earned my B.Sc. in Computer Engineering at Iran University of Science and Technology (IUST), where I worked on Persian NLP and co-authored a mental-health chatbot paper published at AbjadNLP 2025. I've also served as a teaching assistant for courses from Stochastic Processes and Machine Learning to NLP and Computer Architecture, and I build applied healthcare-AI tools such as DentalMind.
Research
Proceedings of the 1st Workshop on NLP for Languages Using Arabic Script (AbjadNLP @ COLING 2025), pp. 64–77, ACL
Survey manuscript — backdoor attacks and defenses for large language models
Open source
A “neuro-statistical” watchdog that flags backdoored LoRA adapters without knowing the trigger, using the suspect model’s confidence and semantic-consistency signals and a multi-head verdict: clean, suspicious, or backdoor detected.
Extends the BAIT (S&P 2025) LLM backdoor scanner with sparsemax candidate selection and a controlled zoo of benign, standard, and evasive backdoored LoRA models to locate exactly where the scanner fails while the backdoor still fires.
A backdoor detection system for LLMs based on attention analysis, introducing a “self-centeredness” score to catch sophisticated backdoors that evade confidence-only detectors.
Trustworthy AI for dental diagnostics: per-tooth radiograph analysis designed to support, not replace, a clinician’s judgment. Python, PyTorch, computer vision, FastAPI and React.
An end-to-end PyTorch reproduction of Devign (NeurIPS 2019) on the authors’ released data — composite AST/CFG/DFG graphs, gated graph layers and the Conv readout — plus a commit-disjoint leakage audit of the benchmark.
Writing
I reproduced Devign (NeurIPS 2019) from scratch on the authors' own data. It lands 11 points below the paper, two-thirds of its test set leaks through shared commits, and a plain CNN beats it.
ReadA review and implementation of MECP-GAP, a graph-neural-network approach to partitioning 5G base stations into MEC service regions that minimizes handover cost under load-balancing constraints.
ReadTraining a small distilled model to recognize the statistical fingerprint of a backdoored language model, instead of searching for a known trigger word.
ReadGallery
Lab life, campus, friends, mountains — and two very good dogs.
Contact
I'm always happy to talk about research collaborations in AI safety and trustworthy AI, open-source work, and interesting machine-learning problems. The fastest way to reach me is by email.