I'm an Assistant Professor at Carnegie Mellon University, jointly appointed in the Language Technologies Institute (LTI) and the Engineering & Public Policy (EPP) Department, with a courtesy appointment in the Software and Societal Systems Department (S3D), and a core member of CyLab (yes, I'm collecting them like Pokémon cards haha). I lead the Looni Lab.
I'm also a Founding Member of Technical Staff at humans&.
My research interests are privacy (particularly contextual integrity and information flow norms), natural language processing, AI for science (particularly chemistry and drug discovery), LLM reasoning, and the societal implications of ML. I explore the interplay between data, its influence on models, and the expectations of the people who regulate and use these models. My work has been recognized by the NCWIT Collegiate Award and the Rising Star in Adversarial ML Award.
Recruiting & collaborations: If you are interested in working with me, please fill out this brief form .
✦ Explanation about my name: I used to publish under Fatemeh which is my legal name. But I now go by Niloofar, the Lily flower in Farsi!
✦ My academic Job-market material (Fall 2024): Research statement · Teaching statement · DEI statement · CV · Job-talk slides
✦ Previously: I was a Research Scientist at Meta AI's FAIR Alignment group (May–Nov 2025) working on LLM privacy, security, and learning about chemistry! Prior to that, I was a postdoctoral scholar at University of Washington, advised by Yejin Choi and Yulia Tsvetkov. I received my PhD from UC San Diego, advised by Taylor Berg-Kirkpatrick, and during that time I was also a part-time researcher/intern at Microsoft Research—working with the Privacy in AI, Algorithms, and Semantic Machines teams on differential privacy, model compression, and data synthesis.
News Highlights
New! Blog post: "How Solving Navier–Stokes Turned Into a Privacy Debate". A primer on what the labs' privacy and data-retention policies actually say: training defaults by tier, the feedback and safety-flag carve-outs, hidden exemptions to opt-outs, and what deleting a chat really does.
Three NeurIPS 2026 decisions: our position paper "LLM Privacy Requires a Lifecycle-Wide Approach" was accepted to the Position Paper Track; "The AI Observatory: A Public Measure of Real-World AI Use" (with Anka Reuel, Shayne Longpre, and colleagues) was accepted to the Evaluations & Datasets Track; and "Reinforcement Learning on Benign Facts Amplifies Leakage of Memorized Private Data" (Renfei Zhang) will appear at the TAE Workshop.
I went on BBC a couple of times this month to talk about the Hugging Face incident and Jacob Coxon's resignation from Anthropic. Watch the interviews (first, second) and read my thoughts on AI safety (in Farsi, with an English transcript).
"Reinforcement Learning Improves Traversal of Parametric Knowledge in LLMs" (Renfei Zhang, M. Kaniselvan) was accepted as an Oral at EMNLP 2026!
Gave the CMU School of Computer Science New Faculty Lightning Talk (Aug 2026), "Learning the Tails!", a 10-minute intro to the Looni Lab: recording and slides.
Our benchmarks are being used to evaluate frontier models: Anthropic's "Automated researchers can reliably mitigate alignment failures" (Aug 2026) uses ConfAIde as a privacy-violation benchmark, and Meta's Muse Spark Safety & Preparedness Report uses CIMemories for its contextual-privacy evaluation.
Featured in MIT Technology Review (Aug 2026): "We still don't know how people are really using AI" — coverage of the AI Observatory, our public measure of real-world AI use built from 24,521 consented user–AI conversations (85,633 turns) across 52 models, co-led with Anka Reuel and Shayne Longpre.
New! Blog post: "Keep Paying the Humans: Notes from an Anti-Debate" — my write-up of the ICML 2026 anti-debate with Ludwig Schmidt on whether to keep funding a human research workforce once AI surpasses us at research.
Heading to ICML 2026 (Seoul, Jul): co-organizing the 2nd MemFM workshop, an oral on "A Science of AI Must Study Learning Dynamics" (main conference) and "Alignment Whack-a-Mole" (MemFM), a debate with Ludwig Schmidt (my write-up), and panels at the WiML Symposium, KAIST@ICML, and the Global AI Frontiers Symposium.
CMU News (Jul 2026): "Carnegie Mellon Researchers Lead Three DOE Genesis Mission Awards" — I'm part of the Brookhaven National Laboratory–led project on multi-agent hypothesis generation from multi-modal scientific data, with Max Simchowitz and Aviral Kumar.
Accepted into the Anthropic AI for Science Program (Jul 2026), supporting test-time exploration and discovery with Claude on SMDD-Bench (with Aviral Kumar, Amir Barati Farimani, and Kevin Han), and received a Google TPU Research Cloud (TRC) Cloud TPU compute grant.
Summer 2026 talks: co-taught the tutorial "Theory of Mind and Application in Educational Context" at the BEA Workshop at ACL 2026 (slides for my segment on privacy and ToM in education); spoke on "Privacy Governance in the Open-Source AI Ecosystem" at the Governing Generative AI as Knowledge Commons workshop in Pittsburgh (slides); and gave "What Should an Agent Remember, Reveal, and Redact?" at the University of Cagliari's Machine Learning Security Seminar (recording).
Received an OpenAI Mental Health Research Grant (co-PI with Adam Perer, CMU) to study the privacy and safety of mental-health AI systems.
New paper "Boundary-targeted Membership Inference Attacks on Safety Classifiers": safety classifiers leak the most exactly where mental-health queries sit, at the decision boundary.
SMDD-Bench: Can LLMs Solve Real-World Small Molecule Drug Design Tasks? Agents reward-hack the oracles instead of showing molecular intuition. See the leaderboard.
"Alignment Whack-a-Mole: Finetuning Activates Verbatim Recall of Copyrighted Books in LLMs" (ICML 2026 MemFM oral).
Gave a keynote at the Simons Institute (UC Berkeley) Workshop on Trust in Decentralized Systems (Mar 2026): "2026 Is the New 2016, but Make It Privacy."
New! Blog post: "A Tiny Experiment on Treating PCOS Through Metabolism, Not Hormones" — a year of notes on metabolism, hormones, food noise, and five months on a low-dose GLP-1.
New! Blog post: "From Black and White to Gray: Redefining Privacy for Language" — why privacy is contextual, why scaling won't fix it, and the research behind ConfAIde and CIMemories.
New! I joined humans& as a Founding Member of Technical Staff! Read about why I joined and what about CMU.
Check out my Writing/Blog section — latest post: "Surviving (and Thriving on) the Academic Job Market" — tips on interviews, health, and habits that kept me sane during the job search.
Appeared on The Information Bottleneck podcast (Jan 2026) with Ravid Shwartz-Ziv & Allen Roush: discussed the future of generative AI, how it's reshaping creative work, accelerating scientific discovery, and the ethical frontier of AI.
Gave a talk at the FAR AI San Diego Alignment Workshop at NeurIPS (Dec 2025): "What Does It Mean for Agentic AI to Preserve Privacy?"
CIMemories: A Compositional Benchmark for Contextual Integrity of Persistent Memory in LLMs — Check out the dataset on Hugging Face!
Featured in Science News Explores (Nov 2025): "5 things to remember when talking to a chatbot" — on AI privacy risks and how chatbots handle personal information.
Our write-up "Privacy Is Not Just Memorization" with Tianshi Li is now available (PDF)! Featured in Help Net Security (Oct 2025).
Gave a keynote at CAMLIS 2025 (Oct 2025): "What Does It Mean for Agentic AI to Preserve Privacy? Mapping the New Data Sinks and Leaks" — Video, slides, and reading list
Quoted in the Washington Post (Aug 2025) on AI hype, evaluation metrics, and how people judge AI capabilities.
Gave a keynote at the L2M2 (Large Language Model Memorization) workshop at ACL (Aug 2025): "Emergent Misalignment Through the Lens of Non-verbatim Memorization"
Gave a keynote at the LLMSec workshop at ACL (Aug 2025): "What does it mean for an AI agent to preserve privacy?" — slides
Our position paper "Machine Unlearning Doesn't Do What You Think: Lessons for Generative AI Policy, Research, and Practice" accepted as Oral at NeurIPS 2025!
Appeared on the Jay Shah Podcast (Feb 2025): "Differential Privacy, Creativity & Future of AI Research in the LLM Era"
Gave an invited keynote at NeurIPS 2024 Red Teaming GenAI workshop (Dec 2024) on A False Sense of Privacy: Semantic Leakage and Non-literal Copying in LLMs — recording (jump to 04:50:00).
Appeared on the Thesis Review podcast with Sean Welleck where I talked about my work on Auditing and Mitigating Safety Risks in Large Language Models.
Selected Publications
For the full list, please refer to my Google Scholar page.
-
Position: LLM Privacy Requires a Lifecycle-Wide Approach
NeurIPS 2026 Position Paper Track
-
The AI Observatory: A Public Measure of Real-World AI Use
NeurIPS 2026 Evaluations & Datasets Track — covered in MIT Technology Review
A. Reuel, S. Longpre, ..., N. Mireshghallah, ...
-
Reinforcement Learning Improves Traversal of Parametric Knowledge in LLMs
EMNLP 2026 (Oral)
R. Zhang, M. Kaniselvan, N. Mireshghallah
-
Reinforcement Learning on Benign Facts Amplifies Leakage of Memorized Private Data
TAE Workshop at NeurIPS 2026
R. Zhang, N. Mireshghallah
-
Et Tu, Brute? Economic Misalignment in Personal AI Agents
Preprint 2026
A. Priyanshu, S. Vijay, B. Jabarian, N. Mireshghallah
-
Leaky Language Models: Stealing Architecture and Inference Optimizations via Per-Token Timing
Preprint 2026
S. Majidi, N. Mireshghallah, K. Taram
-
Boundary-targeted Membership Inference Attacks on Safety Classifiers
Preprint 2026
A. Hughes, A. Goldberg, P. Jha, A. Perer, N. Aletras, N. Mireshghallah
-
SMDD-Bench: Can LLMs Solve Real-World Small Molecule Drug Design Tasks?
ICML 2026 GenBio Workshop — Leaderboard
K. Han, R. Zhang, K. Wei, H. Mahdavi, N. Mireshghallah, A. Barati Farimani
-
ICML 2026 MemFM Workshop (Oral)
X. Liu, N. Mireshghallah, J. C. Ginsburg, T. Chakrabarty
-
Position: Don't Just "Fix it in Post": A Science of AI Must Study Learning Dynamics
ICML 2026 (Oral)
S. Biderman, M. A. Khan, N. Mireshghallah, C. Arnett, F. Barez, N. Saphra
-
Preprint 2026
K. Monteiro, M. Park, A. Ioffrida, A. Sanna, H.-P. Lee, N. Mireshghallah, Y. Wang, S. Das
-
Quantifying the Effect of Test Set Contamination on Generative Evaluations
Preprint 2026
R. Schaeffer, J. Kazdan, ..., N. Mireshghallah, S. Koyejo
-
Learning to Reason in 13 Parameters
Preprint 2026
J. X. Morris, N. Mireshghallah, M. Ibrahim, S. Mahloujifar
-
Memorization Dynamics in Knowledge Distillation for Language Models
Preprint 2026
J. Borkar, K. Chadha, N. Mireshghallah, Y. Zhang, I.-E. Veliche, A. Mitra, D. A. Smith, Z. Xu, D. Garcia-Olano
-
NeurIPS 2025 (Oral Presentation)
A. F. Cooper, ..., N. Mireshghallah, ..., K. Lee
-
CIMemories: A Compositional Benchmark for Contextual Integrity of Persistent Memory in LLMs
ICLR 2026 — Dataset on HuggingFace
N. Mireshghallah, et al.
-
PRIVASIS: Synthesizing the Largest "Public" Private Dataset from Scratch
ICML 2026
H. Kim*, N. Mireshghallah*, M. Duan, R. Xin, S. S. Li, J. Jung, D. Acuna, Q. Pang, H. Xiao, G. E. Suh, S. Oh, Y. Tsvetkov, P. W. Koh, Y. Choi
-
Operationalizing Data Minimization for Privacy-Preserving LLM Prompting
ICLR 2026
J. Zhou, N. Mireshghallah, T. Li
-
Spectrum Tuning: Post-Training for Distributional Coverage and In-Context Steerability
ICLR 2026
T. Sorensen, B. Newman, J. Moore, C. Y. Park, J. Fisher, N. Mireshghallah, L. Jiang, Y. Choi
-
Can Large Language Models Really Recognize Your Name?
FAccT 2026
D. Pham, P. Kairouz, N. Mireshghallah, E. Bagdasarian, C. M. Pham, A. Houmansadr
-
Theory of Mind and Application in Educational Context
BEA Workshop at ACL 2026 (tutorial) — Slides: Privacy and ToM in the Context of Education
E. Farhana, M. Zainab, Q. Wang, N. Mireshghallah, R. van der Meulen, M. J. van Duijn
-
RefGrader: Automated Grading of Mathematical Competition Proofs using Agentic Workflows
NeurIPS 2025 Workshop MATH-AI
H. Mahdavi, N. Mireshghallah, ..., V. Honavar
-
A False Sense of Privacy: Evaluating Textual Data Sanitization Beyond Surface-level Privacy Leakage
SaTML 2026
R. Xin*, N. Mireshghallah*, S. S. Li, M. Duan, H. Kim, Y. Choi, Y. Tsvetkov, S. Oh, P. W. Koh
-
ICLR 2025 (Oral Presentation)
X. Lu, M. Sclar, S. Hallinan, N. Mireshghallah, J. Liu, S. Han, A. Ettinger, L. Jiang, K. Chandu, N. Dziri, Y. Choi
-
Exploring the Limits of Strong Membership Inference Attacks on Large Language Models
NeurIPS 2025
J. Hayes, ..., N. Mireshghallah, ..., A. F. Cooper
-
Position: Privacy Is Not Just Memorization!
Write-up, 2025 — featured in Help Net Security
N. Mireshghallah, T. Li
-
PPML Workshop at CRYPTO 2025
Y. Bae, M. Kim, J. Lee, S. Kim, J. Kim, Y. Choi, N. Mireshghallah
-
HAICOSYSTEM: An Ecosystem for Sandboxing Safety Risks in Human–AI Interactions
COLM 2025
X. Zhou, ..., N. Mireshghallah, ..., M. Sap
-
ParaPO: Aligning Language Models to Reduce Verbatim Reproduction of Pre-training Data
COLM 2025
T. Chen, F. Brahman, J. Liu, N. Mireshghallah, W. Shi, P. W. Koh, L. Zettlemoyer, H. Hajishirzi
-
The Surprising Effectiveness of Membership Inference with Simple N-Gram Coverage
COLM 2025
S. Hallinan, J. Jung, M. Sclar, X. Lu, A. Ravichander, S. Ramnath, Y. Choi, S. P. Karimireddy, N. Mireshghallah, X. Ren
-
Privacy Ripple Effects from Adding or Removing Personal Information in Language Model Training
Findings of ACL 2025 — cited in US Congressional testimony on AI chatbot data privacy
J. Borkar, M. Jagielski, K. Lee, N. Mireshghallah, D. A. Smith, C. A. Choquette-Choo
-
Alpaca against Vicuna: Using LLMs to Uncover Memorization of LLMs
NAACL 2025
A. Kassem*, O. Mahmoud*, N. Mireshghallah*, H. Kim, Y. Tsvetkov, Y. Choi, S. Saad, S. Rana
-
Differentially Private Learning Needs Better Model Initialization and Self-Distillation
NAACL 2025 (Oral Presentation)
I. Ngong, J. Near, N. Mireshghallah
-
Information-Guided Identification of Training Data Imprint in (Proprietary) Large Language Models
NAACL 2025 (Honorable Mention Candidate, Oral Presentation)
A. Ravichander, ..., N. Mireshghallah, ..., Y. Choi
-
WildTeaming at Scale: From In-the-Wild Jailbreaks to (Adversarially) Safer Language Models
NeurIPS 2024
L. Jiang, K. Rao, S. Han, A. Ettinger, F. Brahman, S. Kumar, N. Mireshghallah, X. Lu, M. Sap, Y. Choi, N. Dziri
-
EMNLP 2024
T. Chen, N. Mireshghallah*, A. Asai*, S. Min, J. Grimmelmann, Y. Choi, H. Hajishirzi, L. Zettlemoyer, P. W. Koh
-
Trust No Bot: Discovering Personal Disclosures in Human-LLM Conversations in the Wild
COLM 2024
N. Mireshghallah*, M. Antoniak*, Y. More*, Y. Choi, G. Farnadi
-
Do membership inference attacks work on large language models?
COLM 2024
M. Duan, A. Suri, N. Mireshghallah, S. Min, W. Shi, L. Zettlemoyer, Y. Tsvetkov, Y. Choi, D. Evans, H. Hajishirzi
-
Machine Unlearning Doesn't Do What You Think
Extended Abstract at GenLaw 2024
K. Lee, A. F. Cooper, C. A. Choquette-Choo, K. Liu, M. Jagielski, N. Mireshghallah, L. Ahmed, J. Grimmelmann, D. Bau, C. De Sa, et al.
-
A Roadmap to Pluralistic Alignment
ICML 2024
T. Sorensen, J. Moore, J. Fisher, M. Gordon, N. Mireshghallah, C. M. Rytting, A. Ye, L. Jiang, X. Lu, N. Dziri, T. Althoff, Y. Choi
-
ICLR 2024 (Spotlight)
N. Mireshghallah*, H. Kim*, X. Zhou, Y. Tsvetkov, M. Sap, R. Shokri, Y. Choi
-
Privacy-preserving in-context learning with differentially private few-shot generation
ICLR 2024
X. Tang, R. Shin, H. A. Inan, A. Manoel, N. Mireshghallah,Z. Lin, S. Gopi, J. Kulkarni, R. Sim
-
Smaller Language Models are Better Black-box Machine-Generated Text Detectors
EACL 2024
N. Mireshghallah, J. Mattern, S. Gao, R. Shokri, T. Berg-Kirkpatrick
-
Privacy-Preserving Domain Adaptation of Semantic Parsers
ACL 2023
N. Mireshghallah, R. Shin, Y. Su, T. Hashimoto, J. Eisner
-
Non-Parametric Temporal Adaptation for Social Media Topic Classification
EMNLP 2023
N. Mireshghallah*, N. Vogler*, J. He, O. Florez, A. El-Kishky, T. Berg-Kirkpatrick
-
A Block Metropolis-Hastings Sampler for Controllable Energy-based Text Generation
CoNLL 2023
J. Forristal, N. Mireshghallah, G. Durrett, T. Berg-Kirkpatrick
-
Membership Inference Attacks against Language Models via Neighbourhood Comparison
ACL 2023
J. Mattern, N. Mireshghallah, Z. Jin, B. Scholkop, M. Sachan, T. Berg-Kirkpatrick
-
Differentially Private Model Compression
NeurIPS 2022
N. Mireshghallah, A. Backurs, H. A. Inan, L. Wutschitz, J. Kulkarni
-
Memorization in NLP Fine-tuning Methods
EMNLP 2022
N. Mireshghallah, A. Uniyal, T. Wang, D. Evans, T. Berg-Kirkpatrick
-
Quantifying Privacy Risks of Masked Language Models Using Membership Inference Attacks
EMNLP 2022
N. Mireshghallah, K. Goyal, A. Uniyal, T. Berg-Kirkpatrick, R. Shokri
-
NAACL 2022
N. Mireshghallah, V. Shrivastava, M. Shokouhi, T. Berg-Kirkpatrick, R. Sim, D. Dimitriadis
-
What Does it Mean for a Language Model to Preserve Privacy?
FAccT 2022
H. Brown, K. Lee, N. Mireshghallah, R. Shokri, F. Tram'er
-
Mix and Match: Learning-free Controllable Text Generation
ACL 2022
N. Mireshghallah, K. Goyal, T. Berg-Kirkpatrick
-
Style Pooling: Automatic Text Style Obfuscation for Improved Classification Fairness
EMNLP 2021
N. Mireshghallah, T. Berg-Kirkpatrick
-
Privacy Regularization: Joint Privacy-Utility Optimization in Language Models
NAACL 2021
N. Mireshghallah, H. Inan, M. Hasegawa, V. Rühle, T. Berg-Kirkpatrick, R. Sim
-
ICML 2020
A. Elthakeb, P. Pilligundla, N. Mireshghallah, A. Cloninger, H. Esmaeilzadeh
-
Not All Features Are Equal: Discovering Essential Features for Preserving Prediction Privacy
WWW 2021
N. Mireshghallah, M. Taram, A. Jalali, A. T. Elthakeb, D. Tullsen, H. Esmaeilzadeh
-
Shredder: Learning Noise Distributions to Protect Inference Privacy
ASPLOS 2020
N. Mireshghallah, M. Taram, A. Jalali, D. Tullsen, H. Esmaeilzadeh
Invited Talks
Upcoming
-
INFORMS 2026 Annual Meeting, San Francisco
Invited Speaker, Nov. 2026
Session: Large Language Models and Their Interaction with Social Systems
-
Copenhagen NLP Symposium
Keynote, Oct. 2026
-
RE-Data: Workshop on Responsibly Enabling Data for Foundation Models at COLM 2026
Invited Talk, Oct. 2026
Emergent Misalignment Through the Lens of Memorization
-
CMU School of Computer Science, New Faculty Lightning Talks
Lightning Talk, Aug. 2026
Learning the Tails!
-
AI as a Tool for Mathematics, Computer Science, and Machine Learning Workshop at ICML 2026
Debate (with Ludwig Schmidt), Jul. 2026
Should we keep funding a human research workforce once AI surpasses us at research?
-
Panelist, Jul. 2026
Career Transitions in ML
-
Panelist, Jul. 2026
Career Advice for Junior Researchers & Students
-
Global AI Frontiers Symposium, Seoul
Panelist, Jul. 2026
-
Workshop on Innovative Use of NLP for Building Educational Applications (BEA) at ACL 2026
Tutorial Co-instructor, Jul. 2026
Theory of Mind and Application in Educational Context — my segment: Privacy and ToM in the Context of Education
-
2nd Workshop on Memorization and Trustworthy Foundation Models (MemFM) at ICML 2026
Oral, Jul. 2026
Alignment Whack-a-Mole: Finetuning Activates Verbatim Recall of Copyrighted Books in LLMs
-
University of Cagliari (PRA Lab), Italy
Machine Learning Security Seminar Series, Jun. 2026
What Should an Agent Remember, Reveal, and Redact? The New Data Sinks and Leaks of Agentic AI
-
Google
Invited Talk, Jun. 2026
-
Governing Generative AI as Knowledge Commons Workshop, Pittsburgh
Invited Talk, Jun. 2026
Privacy Governance in the Open-Source AI Ecosystem: Memorization, Provenance, and Contextual Integrity
-
Cornell University, AI Seminar
Seminar, May 2026
2026 Is the New 2016, but Make It Privacy: On Federated Memory, Contextual Privacy, and Personalized Agents
Slides (Simons Institute version)
-
Seminar, Apr. 2026
2026 Is the New 2016, but Make It Privacy: On Federated Memory, Contextual Privacy, and Personalized Agents
-
Keynote, Workshop on Trust in Decentralized Systems, Mar. 2026
2026 Is the New 2016, but Make It Privacy: On Federated Memory, Contextual Privacy, and Personalized Agents
-
FAR AI San Diego Alignment Workshop at NeurIPS
Workshop Keynote, Dec. 2025
What Does It Mean for Agentic AI to Preserve Privacy?
-
CMU CyLab Partners Conference 2025
Invited Talk, Oct. 2025
What Does It "Really" Mean for Agentic AI to Preserve Privacy?
-
Keynote, Oct. 2025
What Does It Mean for Agentic AI to Preserve Privacy? Mapping the New Data Sinks and Leaks
-
Cornell Tech Digital Life Seminar
Seminar, Oct. 2025
Contextual Privacy in LLMs: Benchmarking and Mitigating Inference-Time Risks
-
Meta AI / FAIR Alignment Group
Invited Talk, Oct. 2025
What You Should Really Worry About When It Comes to Generative AI and Privacy
-
First Workshop on LLM Security (LLMSec) at ACL 2025
Keynote, Aug. 2025
What Does It Mean for Agentic AI to Preserve Privacy?
-
First Workshop on Large Language Model Memorization (L2M2) at ACL 2025
Keynote, Aug. 2025
Emergent Misalignment Through the Lens of Non-verbatim Memorization
-
Workshop on Collaborative and Federated Agentic Workflows (CFAgentic) at ICML 2025
Invited Talk, July 2025
What Does It Mean for Agentic AI to Preserve Privacy?
-
Fifth Workshop on Trustworthy Natural Language Processing @NAACL 2025 (TrustNLP)
Invited Talk, May 2025
A False Sense of Privacy: Semantic Leakage and Non-literal Copying in LLMs
-
ETH Zurich
Guest Lecture, May 2025
Privacy, Copyright and Data Integrity: The Cascading Implications of Generative AI
-
UC Berkeley School of Information
Invited Talk, Feb. 2025
Privacy, Copyright, and Data Integrity: The Cascading Implications of Generative AI
-
Stanford University (NLP Seminar)
NLP Seminar, Jan. 2025
Privacy, Copyright and Data Integrity: The Cascading Implications of Generative AI
-
University of California, Los Angeles
Guest lecture for CS 269 - Computational Ethics, LLMs and the Future of NLP, Jan. 2025
Privacy, Copyright and Data Integrity: The Cascading Implications of Generative AI
-
NeurIPS Conference (Red Teaming GenAI workshop)
Red Teaming GenAI workshop, Dec. 2024
A False Sense of Privacy: Semantic Leakage and Non-literal Copying in LLMs
Recording (jump to 04:50:00)
-
NeurIPS Conference (PrivacyML Tutorial)
Panelist, Dec. 2024
PrivacyML: Meaningful Privacy-Preserving Machine Learning tutorial
Recording (jump to 01:52:00)
-
Johns Hopkins University
CS Department Seminar, Dec. 2024
Privacy, Copyright and Data Integrity: The Cascading Implications of Generative AI
-
Future of Privacy Forum
Panelist, Nov. 2024
Technologist Roundtable for Policymakers: Key Issues in Privacy and AI
-
University of Utah
Guest lecture for the School of Computing CS 6340/5340 NLP course, Nov. 2024
Can LLMs Keep a Secret?
-
UMass Amherst
NLP Seminar, Oct. 2024
Membership Inference Attacks and Contextual Integrity for Language
-
Northeastern University
Khoury College of Computer Sciences Security Seminar, Oct. 2024
Membership Inference Attacks and Contextual Integrity for Language
-
Stanford Research Institute (SRI) International
Computational Cybersecurity in Compromised Environments (C3E) workshop, Sep. 2024
Can LLMs keep a secret? Testing privacy implications of Language Models via Contextual Integrity
-
LinkedIn Research
Privacy Tech Talk, Sep. 2024
Can LLMs keep a secret? Testing privacy implications of Language Models via Contextual Integrity
-
National Academies (NASEM)
Forum on Cyber Resilience, Aug. 2024
Oversharing with LLMs is underrated: the curious case of personal disclosures in human-LLM conversations
-
ML Collective
DLCT reading group, Aug. 2024
Privacy in LLMs: Understanding what data is imprinted in LMs and how it might surface!
-
Carnegie Mellon University
Invited Talk, Jun. 2024
Alpaca against Vicuna: Using LLMs to Uncover Memorization of LLMs
-
Generative AI and Law workshop, Washington DC
Invited Talk, Apr. 2024
What is differential privacy? And what is it not?
-
Meta AI Research
Invited Talk, Apr. 2024
Membership Inference Attacks and Contextual Integrity for Language
-
Georgia Institute of Technology
Guest lecture for the School of Interactive Computing, Apr. 2024
Safety in LLMs: Privacy and Memorization
-
University of Washington
Guest lecture for CSE 484 and 582 courses on Computer Security and Ethics in AI, Apr. 2024
Safety in LLMs: Privacy and Memorization
-
Carnegie Mellon University
Guest lecture for LTI 11-830 course on Computational Ethics in NLP, Mar. 2024
Safety in LLMs: Privacy and Memorization
-
Simons Collaboration
TOC4Fairness Seminar, Mar. 2024
Membership Inference Attacks and Contextual Integrity for Language
-
University of California, Santa Barbara
NLP Seminar Invited Talk, Mar. 2024
Can LLMs Keep a Secret? Testing Privacy Implications of LLMs
-
University of California, Los Angeles
NLP Seminar Invited Talk, Mar. 2024
Can LLMs Keep a Secret? Testing Privacy Implications of LLMs
-
University of Texas at Austin
Guest lecture for LIN 393 course on Social Applications and Impact of NLP, Feb. 2024
Can LLMs Keep a Secret? Testing Privacy Implications of LLMs
-
Google Brain
Google Tech Talk, Feb. 2024
Can LLMs Keep a Secret? Testing Privacy Implications of LLMs
-
University of Washington
Allen School Colloquium, Jan. 2024
Can LLMs Keep a Secret? Testing Privacy Implications of LLMs
-
University of Washington
eScience Institute Seminars, Nov. 2023
Privacy Auditing and Protection in Large Language Model
-
CISPA Helmholtz Center for Security
Invited Talk, Sep. 2023
What does privacy-preserving NLP entail?
-
Max Planck Institute for Software Systems
Next 10 in AI Series, Sep. 2023
Auditing and Mitigating Safety Risks in LLMs
-
Mila / McGill University
Invited Talk, May 2023
Privacy Auditing and Protection in Large Language Models
-
EACL 2023
Tutorial co-instruction, May 2023
Private NLP: Federated Learning and Privacy Regularization
-
LLM Interfaces Workshop and Hackathon
Invited Talk, Apr. 2023
Learning-free Controllable Text Generation
-
University of Washington
Invited Talk, Apr. 2023
Auditing and Mitigating Safety Risks in Large Language Models
-
NDSS Conference
Keynote talk for EthiCS workshop, Feb. 2023
How much can we trust large language models?
-
Google
Federated Learning Seminar, Feb. 2023
Privacy Auditing and Protection in Large Language Models
-
University of Texas Austin
Invited Talk, Oct. 2022
How much can we trust large language models?
-
Johns Hopkins University
Guest lecture for CS 601.670 course on Artificial Agents, Sep. 2022
Mix and Match: Learning-free Controllable Text Generation
-
KDD Conference
Adversarial ML workshop, Aug. 2022
How much can we trust large language models?
-
Microsoft Research Cambridge
Invited Talk, Mar. 2022
What Does it Mean for a Language Model to Preserve Privacy?
-
University of Maine
Guest lecture for COS435/535 course on Information Privacy Engineering, Dec. 2021
Improving Attribute Privacy and Fairness for Natural Language Processing
-
National University of Singapore
Invited Talk, Nov. 2021
Style Pooling: Automatic Text Style Obfuscation for Fairness
-
Big Science for Large Language Models
Invited Panelist, Oct. 2021
Privacy-Preserving Natural Language Processing
-
Research Society MIT Manipal
Cognizance Event Invited Talk, Jul. 2021
Privacy and Interpretability of DNN Inference
-
Alan Turing Institute
Privacy and Security in ML Seminars, Jun. 2021
Low-overhead Techniques for Privacy and Fairness of DNNs
-
Split Learning Workshop
Invited Talk, Mar. 2021
Shredder: Learning Noise Distributions to Protect Inference Privacy
-
University of Massachusetts Amherst
Machine Learning and Friends Lunch, Oct. 2020
Privacy and Fairness in DNN Inference
-
OpenMined Privacy Conference
Invited Talk, Sep. 2020
Privacy-Preserving Natural Language Processing
-
Microsoft Research AI
Breakthroughs Workshop, Sep. 2020
Private Text Generation through Regularization
Awards and Honors
Delta Institute Scientific Fellow, 2026
Anthropic AI for Science Program, 2026
Google TPU Research Cloud (TRC) Compute Grant, 2026
Prime Intellect Academic Research Compute Grant, 2026
Greenwall Foundation Faculty Scholars Program in Bioethics, Finalist, 2026
Tinker Academic Research Compute Grant, 2025
Modal Academic Research Compute Grant, 2025
Momental Foundation Mistletoe Research Fellowship (MRF) Finalist, 2023
Rising Star in Adversarial Machine Learning (AdvML) Award Winner, 2022. AdvML Workshop
Rising Stars in EECS, 2022. Event Page
UCSD CSE Excellence in Leadership and Service Award Winner, 2022
FAccT Doctoral Consortium, 2022. FAccT 2022
Qualcomm Innovation Fellowship Finalist, 2021. Fellowship Page
NCWIT (National Center for Women & IT) Collegiate Award Winner, 2020. NCWIT Awards
National University Entrance Exam in Math, 2014. Ranked 249th of 223,000
National University Entrance Exam in Foreign Languages, 2014. Ranked 57th of 119,000
National Organization for Exceptional Talents (NODET), 2008. Admitted, ~2% Acceptance Rate
Research Support
Grants, gifts, and compute programs supporting the Looni Lab. We are grateful to all of our sponsors.
DOE Genesis Mission — collaboration with Brookhaven National Laboratory on multi-agent hypothesis generation from multi-modal scientific data (co-PI), 2026
OpenAI Mental Health Research Grant (lead PI; co-PI Adam Perer), 2026. We thank OpenAI for support for this work from the Mental Health and AI grant program. Funding does not imply endorsement of these results by OpenAI.
Anthropic AI for Science Program — Claude API credits for test-time exploration and discovery on SMDD-Bench (lead PI), 2026
Foresight Institute — AI for Science & Safety gift (lead PI), 2026
Google TPU Research Cloud (TRC) — Cloud TPU compute grant, 2026. Research supported with Cloud TPUs from Google's TPU Research Cloud (TRC).
Google Gemini Academic Research Program — compute grant, 2026
Academic compute grants: Thinking Machines (Tinker), Modal, Prime Intellect, Lambda, and River, 2025–2026
Featured Press & Media
The Information Bottleneck podcast (Jan 2026) — on the future of generative AI with Ravid Shwartz-Ziv & Allen Roush
Jay Shah Podcast (Feb 2025) — on Differential Privacy, Creativity & Future of AI Research
Thesis Review podcast with Sean Welleck — on Auditing and Mitigating Safety Risks in LLMs
Should I do a postdoc — guest video on Sasha Rush's channel + blog post
BBC Persian, ۶۰ دقیقه (60 Minutes) - live segment on AI risks and society (Sep 2026; segment starts at 47:00) · second interview
MIT Technology Review - We still don't know how people are really using AI (Aug 2026; on the AI Observatory)
Recent Co-organized Workshops & Service
[for full list check my CV]Publicity Chair, ACL 2027
Senior Program Committee, AAAI 2027 (Alignment Track)
2nd Workshop on Memorization and Trustworthy Foundation Models (MemFM) @ICML 2026 (Co-organizer)
Workshop on Foundations of Deep Generative Models (FoGen) @ICML 2026 (Co-organizer)
Advisory Committee, 2nd Workshop on Technical AI Governance Research (TAIGR) @ICML 2026
Memorization and Trustworthy Foundation Models Workshop @ICML 2025 (Co-organizer)
Area Chair for NeurIPS 2026 and COLM 2025 & 2026
Workshop on Technical AI Governance (TAIG) @ICML 2025 (Panelist)
Workshop on Collaborative and Federated Agentic Workflows (CFAgentic) @ICML 2025 (Panelist)
Privacy Session Chair at SAGAI Workshop @IEEE S&P 2025
Industry Research Experience
-
Microsoft Semantic Machines
Fall 2022-Fall 2023 (Part-time), Summer 2022 (Intern)
Mentors: Richard Shin, Yu Su, Tatsunori Hashimoto, Jason Eisner
-
Microsoft Research, Algorithms Group, Redmond Lab
Winter 2022 (Intern)
Mentors: Sergey Yekhanin, Arturs Backurs
-
Microsoft Research, Language, Learning and Privacy Group, Redmond Lab
Summer 2021 (Intern), Summer 2020 (Intern)
Mentors: Dimitrios Dimitriadis, Robert Sim
-
Western Digital Co. Research and Development
Summer 2019 (Intern)
Mentor: Anand Kulkarni
Diversity, Inclusion & Mentorship
Mentor for Women in Machine Learning (WiML) Workshop at NeurIPS 2025
Panelist at CMU School of Computer Science Panel: Navigating the Academic Job Market (2025)
Mentor on the 'How to broadcast your research to a wider audience?' panel at ACL Mentorship Program -- 2025
Mentor for the mentorship program at WiML event in NeurIPS 2024
D&I chair at NAACL 2025
Widening NLP (WiNLP) co-chair, 2022–2024
Socio-cultural D&I chair at NAACL 2022
Mentor for the Graduate Women in Computing (GradWIC) at UCSD
Mentor for the UC San Diego Women Organization for Research Mentoring (WORM) in STEM
Co-leader for the "Feminist Perspectives for Machine Learning & Computer Vision" Break-out session at the Women in Machine Learning (WiML) 2020 Un-workshop Held at ICML 2020
Mentor for the USENIX Security 2020 Undergraduate Mentorship Program
Volunteer at the Women in Machine Learning 2019 Workshop Held at NeurIPS 2019
Invited Speaker at the Women in Machine Learning and Data Science (WiMLDS) NeurIPS 2019 Meetup
Mentor for the UCSD CSE Early Research Scholars Program (CSE-ERSP) in 2018
Professional Services
Area Chair, NeurIPS 2026
Area Chair, COLM 2025 & 2026
Senior Program Committee, AAAI 2027 (Alignment Track)
Publicity Chair, ACL 2027
Area Chair, NAACL 2025
Area Chair, EMNLP 2024
Program Committee Member, ACM CCS 2024 & 2025
Diversity & Inclusion Co-chair, NAACL 2022 & 2025
Co-chair, Widening NLP (WiNLP), 2022–2024
Reviewer, ICLR 2021–2025 (Outstanding Reviewer Award, 2021)
Reviewer, NeurIPS and ICML, 2020–2024
Reviewer, AISTATS 2025; Reviewer for ICLR & ICML Workshop Proposals, 2025
Reviewer, COLM, EACL, and CHI, 2024
Reviewer, AAAI and FAccT, 2023–2024
Reviewer, TACL (2020–present), TMLR (2022–present), and IEEE Security & Privacy Magazine (2021–present)
Reviewer, NeurIPS Workshop Proposals and INLG, 2023
Ethics PC Member, EMNLP 2022; Reviewer, NAACL SRW 2022; PC Member, AFCP & TSRML Workshops at NeurIPS 2022
Shadow PC Member, IEEE S&P 2021; Artifact Evaluation PC, USENIX Security 2021; Reviewer, CCS Posters 2021
Artifact Evaluation PC, ASPLOS 2020; PC Member, LatinX in AI & WHI Workshops at ICML 2020 and MLArchSys at ISCA 2020; Security & Privacy Committee Member and Session Chair, Grace Hopper Celebration 2020
Reviewer, ACM TACO and IEEE TC Journals, 2020–present
Books I Like!
Range: Why Generalists Triumph in a Specialized World by D. Epstein
Messy: The Power of Disorder to Transform Our Lives by T. Harford
Small Is Beautiful: Economics As If People Mattered by E. F. Schumacher
Quarter-life by Satya Doyle Byock
The Body Keeps the Score by Bessel van der Kolk
36 Views of Mount Fuji by Cathy Davidson
Indistractable by Nir Eyal
Sapiens: A Brief History of Humankind by Yuval Noah Harari
The Martian by Andy Weir
The Solitaire Mystery by Jostein Gaarder
The Orange Girl by Jostein Gaarder
Life is Short: A Letter to St Augustine by Jostein Gaarder
The Alchemist by Paulo Coelho