Namaste, I’m Himanshu Beniwal 🙏🏻
Postdoctoral Researcher at ScaDS.AI, TU Dresden 🇩🇪 — teaching machines to read the world’s languages safely, truthfully, and interpretably.
- 21Publications
- 314Citations
- 5Fellowships & Awards
- 16Teaching Assignments
- 5Live Projects
About Me 🫡
I am currently a Postdoctoral Researcher at ScaDS.AI, Technische Universität Dresden — TU Dresden 🇩🇪, advised by Prof. Michael Färber. My research focuses on multilingual NLP, AI safety, and mechanistic interpretability, with a particular emphasis on building reliable, truthful, and safe LLMs.
I completed my Ph.D. at IIT Gandhinagar 🇮🇳, advised by Prof. Mayank Singh. My research broadly lay at the intersection of Robust and Interpretable NLP, where I focused on assessing factuality, toxicity, and safety in large language models. I was particularly interested in building interpretable and explainable NLP systems that are reliable across diverse languages and cultures.
During my Ph.D., I was a Visiting PhD Intern at the University of Virginia 🇺🇸, working with Prof. Thomas Hartvigsen on multilingual and interpretable content moderation. Together, our works showed that safety systems built for English alone fail to generalize, motivating multilingual safety moderation, cross-lingual detoxification, and interpretable methods to surface the hidden criteria that communities actually use to enforce their norms.
During my Ph.D., I was also a PhD Intern at Microsoft Research India 🇮🇳, working with Dr. Sunayana Sitaram on multilingual evaluation. Together, our works found that current multilingual LLM benchmarks are neither comprehensive nor explanatory, and proposed a Bayesian framework to decompose and diagnose why performance disparities emerge across languages.
Fellowships during PhD: Prime Minister’s Research Fellowship (PMRF), Microsoft Research India PhD Award ‘25, and Overseas Research Fellowship ‘24; and am a recipient of the Fulbright-Nehru Doctoral Fellowship ‘25. Our work on Cross-lingual Model Editing with Prof. Mayank Singh was also supported by a Microsoft Accelerate Foundation Models Research award (Sept 2023).
More about me! 💭
- 🔭 I’m currently working on how an AI reads & understand the language! 🤖
- 📫 Socials: LinkedIn 👨🏼💼, Twitter 🐤, Hugging-Face 🤗
- 📸 Travel Pics: Instagram
- 😄 Fav mathematical equation: The magic of Euler’s Identity; \(e^{i \pi} + 1 = 0\)
- ⚡ Fun fact: Traveling the 🌎 with 🖤 for espresso ☕️ & crazy for 💻.
Experience 🧑🏻🔬
-
Postdoctoral Researcher Technische Universität Dresden · ScaDS.AI Trustworthy and Robust LLMs for Multilingual Natural Language Processing, with Prof. Michael Färber.
-
PhD Research Intern Microsoft Research India Multicultural, multilingual and multimodal evaluation, with Dr. Sunayana Sitaram. Recipient of the Microsoft Research India PhD Award 2025.
-
PhD Research Intern University of Virginia Cross-lingual detoxification in LLMs using model editing, with Prof. Tom Hartvigsen and Dr. Soham Dan. Supported by the Overseas Research Fellowship from IIT Gandhinagar.
-
Research Student — CodeStream Research Group H. N. B. Garhwal University (A Central University) Deep Q-Learning algorithms for self-driving cars, with Prof. Narender Kumar Rawal.
-
Social Networks Analysis Research Intern Indian Institute of Technology Ropar Node sampling techniques and centrality measures in large-scale networks, with Prof. Sudarshan Iyengar.
Research Interests 🤯
News 🔊
- [September 2026] Presenting at ResAI 2026: Resilience and AI Workshop.
- [September 2026] Presenting at the Panel at 4th IÖR Conference “Space & Transformation”.
- [September 2026] Attending the EUTOPIA Impact School 2026.
- [September 2026] Attending the Germany-Poland-Czechia Workshop on Neurosymbolic and Trustworthy AI.
- [September 2026] Presenting a talk at Resilient AI: Securing Large Language Models!
- [August 2026] Our visionary idea paper on “From Universal Knowledge Graphs to Contextual Semantic Contracts” got accepted at ISWC 2026.
- [August 2026] DEPART got accepted at EMNLP ‘26 🇭🇺!
- [July 2026] “The State and Fate of Multilingual, Contextual Evaluation” is accepted at COLM 2026! Check here! 🔥
- [July 2026] Joined Post-Doctorate at ScaDS.AI / Technische Universität Dresden — TU Dresden 🇩🇪 with Prof. Michael Färber. 🎉🔥
- [May 2026] Defended my PhD thesis! 🎉🎉🎉
- [April 2026] 2️⃣ papers got accepted at ACL 2026! (Beyond Monolingual Assumptions and A Survey of Toxicity Mitigation Strategies for mLLMs) at Mains and Findings! 🔥🔥
- [April 2026] Preprint on “The State and Fate of Multilingual, Contextual Evaluation” is out now! Check here! 🔥
- [February 2026] Presented a talk on using AI for social intelligence through Arbiter at the India - AI Impact Summit 2026!! Check here!!
- [February 2026] Our UnityAI-Guard work got accepted to be presented at the Global South Research & Posters Showcase at the Research Symposium on AI and its Impact, as part of India - AI Impact Summit 2026!! ⭐️
- [February 2026] Our TempUN got accepted to be presented at ACM India ARCS 2026!
- [February 2026] Visiting Microsoft Research India in Bangalore, India 🇮🇳 for my PhD internship!! 🎉
- [January 2026] Submitted my PhD thesis! 🎉🎉🎉
- [November 2025] Our paper Comi-lingua got featured in “9 Indian AI Research Papers From 2025 That Deserve More Attention” - Analytics India Mag.
- [September 2025] 4️⃣ papers (Decoding the Rule Book, COMI-LINGUA, UnityAI-Guard, and Char-mander 🔥) are accepted at EMNLP 2025 (1 each at Mains, Findings, Demo, and Workshop)!!!!
- [August 2025] 2️⃣ papers got accepted at COLM 2025! (PolyGuard and Breaking mBad at MELT Workshop!)
- [June 2025] Awarded with the Microsoft Research India PhD Award 2025!!! 🎉🎉🎉
- [June 2025] Recipient of the prestigious Fulbright Nehru Doctoral Fellowship 2025-2026! ⭐️⭐️⭐️
- [April 2025] Talk on “Cross-lingual Backdoors” at Plutous! [Recording] ⭐️
- [March 2025] Attended Advanced Language Processing School (ALPS) 2025 at the beautiful Centre CNRS Paul Langevin, Aussois, France! 🇫🇷
- [March 2025] Talk on “GenAI in HealthCare”, at Google Developer Group - Silver Oak University, Ahmedabad, India 🇮🇳!!
- [March 2025] Attended PMRF Symposium 2025 at IIT Hyderabad, India 🇮🇳!
- [Feb 2025] Welcome the UnityAI-Guard 😵💫 to the world 🌎!
- [Feb 2025] Char-mander 🔥, “A Study of Cross-lingual Backdoor Attacks in Multilingual LLMs”, is now on ArXiv 🕸️!
- [Feb 2025] Attended Pingala Interactions in Computing (PIC) 2025 at the fabulous Mysore campus of Infosys, India! 🇮🇳
- [Jan 2025] Attended Google DeepMind Research Symposium 2025 at Google Office, Bangalore, India! 🇮🇳
- [Sept 2024] 3️⃣ papers (TempUN, PythonSaga, and COMMENTATOR) made it to EMNLP 24’ 🇺🇸! 🏝️
- [Aug 2024] Talk on “Editing Large Language Models”, MilaNLP, Italy! 🇮🇹
- [Aug 2024] Visiting the University of Virginia (🇺🇸) with Prof. Tom Hartvigsen!! Super excited to work on Interpretable + Multilingual NLP! 🤯🤯🤯
- [Aug 2024] Attended ACL 2024 at Bangkok, Thailand! 🇹🇭
- [July 2024] Released the Ganga-1B model! The Ganga-1B model outperforms existing open-source models that support Indian languages, even at sizes of up to 7 billion parameters.
- [June 2024] Gave a talk at ACM Summer School on Generative AI for Text on “Instruction fine-tuning, FLAN-T5, and Quantisation”! 🤯 (Recording)
- [June 2024] Gave a talk at India-ML Reading Group on “Editing LLMs”! 🤩 (Recording)
- [March 2024] XME Made it to EACL 2024, March 17-22, 2024, at St. Julian’s, Malta! 🇲🇹
- [March 2024] Gave a talk on “Editing Large Language Models” at Google Research India, Bangalore, India!! 🇮🇳
- [March 2024] Attended PMRF Symposium 2024, March 3-4, 2024, at IIT Indore, India.
- [Feb 2024] Attended Research Week with Google 2024, Feb 1-3, 2024, at Bengaluru, India.
- [Jan 2024] XME is now on ArXiv 🕸️!
- [Dec 2023] Gandhipedia, a project of National Importance by joint initiative of IIT Kharagpur, IIT Gandhinagar, and NCSM (Kolkata), under the aegis of The Ministry of Culture, Government of India, was launched! 📢📢📢
- [Dec 2023] Served as Publicity Chair for IndoML 2023 at IIT Bombay.
- [Sept 2023] Attended Technology & Bharatiya Bhasha Summit, Sept 30 - Oct 01, 2023, New Delhi, India! 🇮🇳
- [July 2023] Attended Deep Learning and Artificial Intelligence Summer/Winter School 2023 (DLAI7), 17 - 21 July 2023.
- [May 2023] Attended MLSS 2023 at Krakow, Poland! 🇵🇱
- [April 2023] Done with Ph.D. Thesis Proposal Defense! 😊✅
- [Feb 2023] Attended ARCS 2023 and ACM Annual Event 2023 at OIST Bhopal, India.
- [Jan 2023] Attended Research Week with Google 2023 at Bengaluru, India.
- [Jan 2023] Attended CODS-COMAD 2023 at IIT Bombay, India.
- [Dec 2022] Done with PhD Qualifiying Examinations! 😄🕺🏻
- [Oct 2022] Got selected as a Prime Minister’s Research Fellow (PMRF). 📢📢
- [Aug 2022] Attended Oxford Machine Learning Summer School (OxML) 2022.
- [July 2022] Attended Eastern European Machine Learning Summer School (EEML) 2022.
- [Feb 2022] Attended Research Week with Google 2022.
- [Jan 2022] Attended Advanced Language Processing Winter School (ALPS) 2022.
- [August 2021] Started doctoral journey with Prof. Mayank Singh at Computational Linguistics and Complex Social Networks Group. 📢📢📢
Education 👨🏻🎓
-
Indian Institute of Technology Gandhinagar Doctor of Philosophy in Computer Science & Engineering Thesis: Assessing Factuality and Toxicity in Large Language Models. Supervisor: Prof. Mayank Singh.
-
Central University of Punjab Master of Technology in Computer Science & Technology Thesis: Assessing Empathetic Capabilities in Conversational Approaches. Supervisor: Prof. Satwinder Singh.
-
Hemvati Nandan Bahuguna Garhwal University (A Central University) Bachelor of Technology in Computer Science & Engineering Thesis: Autonomous Driving System simulation using Deep Q-Learning in CARLA. Supervisor: Prof. Narender Kumar Rawal.
Live Projects 📽️
MedicaLLM Evaluation Benchmarking mLLMs for maternal health in English, Hindi, and Marathi. Check here
IITGnGPT Classifies AI-generated content into 4 classes across 4 model types from PDFs or text inputs. Try here
UnityAI-Guard Detects toxicity in six low-resource Indian languages — Hindi, Tamil, Telugu, Marathi, Urdu, and Punjabi. Try here
UnityAI-Guard 2.0 Extends toxicity detection to 17 fine-grained categories across six Indian languages — Bengali, Odia, Malayalam, Kannada, Hindi, and Gujarati. Try here
Backdoor Attacks in CV + NLP Demonstrates backdooring in YOLO (a trigger causes person non-detection) and analogous backdoor vulnerabilities across classification, generation, and translation. YOLO · Classification · Generation · Translation
Publications 📚
Citations: 314 Google Scholar profile →
- From Universal Knowledge Graphs to Contextual Semantic Contracts
- DEPART: DEcomposing PARiTy across Multilingual LLMs
- Sycophancy as a Multilingual Alignment Failure: How Safety Degrades Across Languages, Topics, and Models
- Where Does Toxicity Live? Mechanistic Localization and Targeted Suppression in Language Models
- The State and Fate of Multilingual, Contextual Evaluation in the NLP World
- One Instruction Does Not Fit All: How Well Do Embeddings Align Personas and Instructions in Low-Resource Indian Languages?
- A Survey of Toxicity Detection and Mitigation Strategies for Multilingual Language Models
- Beyond Monolingual Assumptions: A Survey of Code-Switched NLP in the Era of Large Language Models
- Decoding the Rule Book: Extracting Hidden Moderation Criteria from Reddit Communities
- COMI-LINGUA: Expert Annotated Large-Scale Dataset for Multitask NLP in Hindi-English Code-Mixing
- UNITYAI-GUARD: Pioneering Toxicity Detection Across Low-Resource Indian Languages
- Char-mander Use mBackdoor! A Study of Cross-lingual Backdoor Attacks in Multilingual LLMs
- Breaking mBad! Supervised Fine-tuning for Cross-Lingual Detoxification
- PolyGuard: A Multilingual Safety Moderation Tool for 17 Languages
- COMMENTATOR: A Code-mixed Multilingual Text Annotation Framework
- PythonSaga: Redefining the Benchmark to Evaluate Code Generating LLMs
- Remember This Event That Year? 🤔 Assessing Temporal Information and Reasoning in Large Language Models
- Cross-lingual Editing in Multilingual Language Models
- Explainable Transformer-based Anomaly Detection for IoT Security
- A survey on near-human conversational agents
- Handwritten Digit Recognition using Machine Learning
Posters & Talks 🔊
- [April 2025] Talk on “Cross-lingual Backdoors” at Plutous! [Recording] ⭐️
- [March 2025] Talk on “GenAI in HealthCare”, at Google Developer Group - Silver Oak University, Ahmedabad, India 🇮🇳!!
- [Sept 2024] Talk on “Editing Large Language Models”, MilaNLP, Italy! 🇮🇹
- [March 2024] Talk on “Editing Large Language Models”, Google Research India, Bangalore, India.
- [March 2024] PMRF Symposium 2024, “Cross-lingual Editing in Multilingual Language Models”, at IIT Indore.
- [Feb 2024] [Research Week with Google 2024] at Google Research India, Bangalore, India, on ‘XME: Cross-lingual Model Editing in LLMs’.
- [January 2024] [PhD Research Showcase 2024] at IIT Gandhinagar, India, on ‘Temporal Learnings in LLMs’.
- [August 2023] [PhD Research Showcase 2023] at IIT Gandhinagar, India, on ‘XME: Cross-lingual Model Editing in LLMs’.
- [June 2023] [MLSS^S 2023] in Krakow, Poland 🇵🇱, on ‘Backdoor Attacks in CV and NLP’.
- [Feb 2023] Talk on “Backdoor Attacks in NLP”, at IISER Bhopal.
See the poster & talk gallery 🤩
Community Experience 👷🏻♂️
- Volunteer: Communications Team at ACL Rolling Review (April 24’ - Present).
- Mentor: Research Mentor at SimPPL (Jan 2024 - Present).
- Member: Web Developer at Research Society (अन्वेषणम्) at IIT Gandhinagar (August 2023 - Present).
- Organizer: IndoML 2023
- Volunteer: IndoML 2022, ACM-IKDD Summer School 2022
- Conference Reviewer: LREC-COLING 2024, EACL CASE 2024, EACL Demo 2024, EAI SaSeIoT 2023, EMNLP 2023, ICTIR 2023, ACL Workshop BigScience 2022, DLSM 2021
- Journal Reviewer: ACI 2022
- Organized 20+ workshops/hackathon events. Pictures 📸
- Beta Reviewer: Coursera
- Mentor: Summer Internship Mentor at RightApprise 2018
- Campus Representative/Ambassador: Google Crowdsource 2019, GeeksforGeeks 2018-19, Internshala 2017-18
- Scholar: Udacity Facebook Scholar 2019, Google India Scholar 2018
Teaching Assistantships 🛳️
- Artificial Intelligence and Natural Language Processing, January 2026 to May 2026.
- IT549: Deep Learning, 45+ students, Jan 2025 to May 2025, with Prof. Arpit Rana.
- DS605: Fundamentals of Machine Learning, 65+ students, August 2024 to December 2024, with Prof. Arpit Rana.
- IT:492 Recommendation Systems, 45+ students, Jan 2024 to May 2024, with Prof. Arpit Rana.
- IT:496 Introduction to Data Mining, July 2023 to December 2023, with Prof. Arpit Rana.
- IT:492 Recommendation Systems, Jan to May 2023, with Prof. Arpit Rana.
- CS:613 Natural Language Processing, August 2025 to December 2025, with Prof. Mayank Singh.
- CS 203: Software Tools & Techniques for AI, January 2025 to May 2025, with Prof. Mayank Singh.
- CS:613 Natural Language Processing, August 2024 to December 2024, with Prof. Mayank Singh.
- [Graduate Teaching Fellow] ES113: Data Centric Computing, Jan 2024 to May 2024, with Prof. Manoj Gupta and Prof. Mayank Singh.
- CS:613 Natural Language Processing, July 2023 to December 2023, with Prof. Mayank Singh.
- ES:432 Databases, Jan to May 2023, with Prof. Mayank Singh.
- ACM-IKDD Summer School on Data Science, July 4th – 16th, 2022.
- ES:432 Databases, Jan to May 2022, with Prof. Mayank Singh.
- ES 102 - Introduction to Computing, Nov to Dec 2021, with Prof. Sairam Swaroop Mallajosyula & Prof. Nipun Batra.
- ES:242 (Data Structure & Algorithms - 1), August to Nov. 2021, with Prof. Manoj Gupta.
Media Coverage 📰
- [Feb 2024] Attended Research Week with Google 2024 at Google Research India, Bangalore, India. Twitter.
- [Dec 2023] Gandhipedia launch! 😀 ETV, ZeeNews, Times of India, The Statesman, and ETV Bharat.
- [Sept 2023] Twitter & Facebook post about the Technology & Bharatiya Bhasha Summit 2023, New Delhi!
- [July 2023] Twitter post about MLSS^S poster presentation.
- [July 2023] Research Capsule Research Showcase at IIT Gandhinagar: LinkedIn, Twitter, Facebook, and Instagram.
- [Nov 2022] PMRF coverage: IITGN News, NDTV News, Careers 360, and others.
Students Mentored 🧑🏻💻
Notebooks 📒
Recommendation Systems
- Introduction to NLP!
- Pretrained Embeddings
- Pretraining and Finetuning.
- User-Collaborative Systems
- Sequential Recommendations
- Sequential RecSystems-2
- Time Series
Introduction to Data Mining
Past Projects 👨🏻💻
Backdoor Attacks in Computer Vision Tasks Himanshu Beniwal, Prof. Shanmuganathan Raman
Explored backdoor attack in MNIST, CIFAR10, MOT, and real-world datasets. Reporting, 99% attack success rate with 0.1% poisoning budget. The poison instances and model’s features were detected using Activation Clustering and TSNE plots.
Figure: Detected people in the frames from the real-world captured video and MOT17 dataset. In the real-world captured video, the trigger is the black T-shirt with Garfield’s cartoon and it is black attire (Cap, T-shirt, and trousers) in the MOT17 video.
Poisoning Attacks in Text Classification and Generation Himanshu Beniwal, Prof. Mayank Singh
Experimented with clean-label and label-flipping attacks in text generation and classification. Achieving 99% ASR with 95% clean-accuracy on SST-2 for classification. Classification models with triggers: ‘Google’, ‘James Bond’, and ‘cf’. Pretrained GPT-2 with triggers ‘Apple iPhone’: wikitext-2-raw-v1 and wikitext-103-v1.
Assessing Empathetic Capabilities in Conversational Approaches Himanshu Beniwal, Prof. Satwinder Singh
To assess the empathetic capabilities in conversational approaches using seq2seq and transformers variations like generative, bi-encoder, poly-encoder, and ranker for empathetic dialogue dataset.
Last updated: August 24, 2026 · Photo gallery 📸