Multilingual NLP · AI Safety · Interpretability

Namaste, I’m Himanshu Beniwal 🙏🏻

Postdoctoral Researcher at ScaDS.AI, TU Dresden 🇩🇪 — teaching machines to read the world’s languages safely, truthfully, and interpretably.

  • 21Publications
  • 314Citations
  • 5Fellowships & Awards
  • 16Teaching Assignments
  • 5Live Projects

About Me 🫡

I am currently a Postdoctoral Researcher at ScaDS.AI, Technische Universität Dresden — TU Dresden 🇩🇪, advised by Prof. Michael Färber. My research focuses on multilingual NLP, AI safety, and mechanistic interpretability, with a particular emphasis on building reliable, truthful, and safe LLMs.

I completed my Ph.D. at IIT Gandhinagar 🇮🇳, advised by Prof. Mayank Singh. My research broadly lay at the intersection of Robust and Interpretable NLP, where I focused on assessing factuality, toxicity, and safety in large language models. I was particularly interested in building interpretable and explainable NLP systems that are reliable across diverse languages and cultures.

During my Ph.D., I was a Visiting PhD Intern at the University of Virginia 🇺🇸, working with Prof. Thomas Hartvigsen on multilingual and interpretable content moderation. Together, our works showed that safety systems built for English alone fail to generalize, motivating multilingual safety moderation, cross-lingual detoxification, and interpretable methods to surface the hidden criteria that communities actually use to enforce their norms.

During my Ph.D., I was also a PhD Intern at Microsoft Research India 🇮🇳, working with Dr. Sunayana Sitaram on multilingual evaluation. Together, our works found that current multilingual LLM benchmarks are neither comprehensive nor explanatory, and proposed a Bayesian framework to decompose and diagnose why performance disparities emerge across languages.

🏅

Fellowships during PhD: Prime Minister’s Research Fellowship (PMRF), Microsoft Research India PhD Award ‘25, and Overseas Research Fellowship ‘24; and am a recipient of the Fulbright-Nehru Doctoral Fellowship ‘25. Our work on Cross-lingual Model Editing with Prof. Mayank Singh was also supported by a Microsoft Accelerate Foundation Models Research award (Sept 2023).

More about me! 💭

  • 🔭 I’m currently working on how an AI reads & understand the language! 🤖
  • 📫 Socials: LinkedIn 👨🏼‍💼, Twitter 🐤, Hugging-Face 🤗
  • 📸 Travel Pics: Instagram
  • 😄 Fav mathematical equation: The magic of Euler’s Identity; \(e^{i \pi} + 1 = 0\)
  • ⚡ Fun fact: Traveling the 🌎 with 🖤 for espresso ☕️ & crazy for 💻.

Experience 🧑🏻‍🔬

Research Interests 🤯

Artificial Intelligence Multilingual NLP Multicultural NLP Multimodal NLP Toxicity Mitigation Stereotypes & Bias Model Editing Temporal Reasoning
Natural Language Processing Robust & Interpretable NLP Secure NLP Model Editing Word Segmentation Conversational AI
Machine Learning Adversarial Attacks Poisoning Attacks Data Poisoning Embedding Poisoning

News 🔊

Education 👨🏻‍🎓

Live Projects 📽️

MedicaLLM Evaluation Benchmarking mLLMs for maternal health in English, Hindi, and Marathi. Healthcare Multilingual Check here

IITGnGPT Classifies AI-generated content into 4 classes across 4 model types from PDFs or text inputs. Detection Demo Try here

UnityAI-Guard Detects toxicity in six low-resource Indian languages — Hindi, Tamil, Telugu, Marathi, Urdu, and Punjabi. Safety 6 languages Try here

UnityAI-Guard 2.0 Extends toxicity detection to 17 fine-grained categories across six Indian languages — Bengali, Odia, Malayalam, Kannada, Hindi, and Gujarati. Safety 17 categories Try here

Backdoor Attacks in CV + NLP Demonstrates backdooring in YOLO (a trigger causes person non-detection) and analogous backdoor vulnerabilities across classification, generation, and translation. Security Poisoning YOLO · Classification · Generation · Translation

Publications 📚

Citations: 314 Google Scholar profile →

  • From Universal Knowledge Graphs to Contextual Semantic ContractsHimanshu Beniwal, Michael FaerberPreprint · May 2026Core Rank: A PDF soon
  • DEPART: DEcomposing PARiTy across Multilingual LLMsManan Uppadhyay, Prashant Kodali, Pranjal Chitale, Reshma Ramaprasad, Himanshu Beniwal, Sunayana SitaramEMNLP 2026 · Findings 🇭🇺Core Rank: A* PDF
  • Sycophancy as a Multilingual Alignment Failure: How Safety Degrades Across Languages, Topics, and ModelsArya Shah, Himanshu Beniwal, Mayank Singh, Chaklam SilpasuwanchaiPreprint · May 2026 PDF
  • Where Does Toxicity Live? Mechanistic Localization and Targeted Suppression in Language ModelsHimanshu Beniwal, Mayank SinghPreprint · May 2026 PDF
  • The State and Fate of Multilingual, Contextual Evaluation in the NLP WorldManan Uppadhyay, Himanshu Beniwal, Prashant Kodali, Sunayana SitaramCOLM 2026 🇺🇸 PDF
  • One Instruction Does Not Fit All: How Well Do Embeddings Align Personas and Instructions in Low-Resource Indian Languages?Arya Shah, Himanshu Beniwal, Mayank SinghArXiv · January 2026 PDF
  • A Survey of Toxicity Detection and Mitigation Strategies for Multilingual Language ModelsSoham Dan, Himanshu Beniwal, Thomas HartvigsenACL 2026 · Findings 🇺🇸Core Rank: A* PDF
  • Beyond Monolingual Assumptions: A Survey of Code-Switched NLP in the Era of Large Language ModelsRajvee Sheth, Samridhi Raj Sinha, Mahavir Patil, Himanshu Beniwal, Mayank SinghACL 2026 · Main 🇺🇸Core Rank: A* PDF
  • Decoding the Rule Book: Extracting Hidden Moderation Criteria from Reddit CommunitiesYoungwoo Kim, Himanshu Beniwal, Steven L. Johnson, Thomas HartvigsenEMNLP 2025 · Main 🇨🇳Core Rank: A* PDF
  • COMI-LINGUA: Expert Annotated Large-Scale Dataset for Multitask NLP in Hindi-English Code-MixingRajvee Sheth, Himanshu Beniwal, Mayank SinghEMNLP 2025 · Findings 🇨🇳Core Rank: A* PDF
  • UNITYAI-GUARD: Pioneering Toxicity Detection Across Low-Resource Indian LanguagesHimanshu Beniwal, Reddybathuni Venkat, Rohit Kumar, Birudugadda Srivibhav, Daksh Jain, Pavan Deekshith Doddi, Eshwar Dhande, Adithya Ananth, Kuldeep, Mayank SinghEMNLP 2025 · Demo 🇨🇳Core Rank: A* PDF Demo
  • Char-mander Use mBackdoor! A Study of Cross-lingual Backdoor Attacks in Multilingual LLMsHimanshu Beniwal, Sailesh Panda, Birudugadda Srivibhav, Mayank SinghBlackboxNLP @ EMNLP 2025 🇨🇳 PDF
  • Breaking mBad! Supervised Fine-tuning for Cross-Lingual DetoxificationHimanshu Beniwal, Youngwoo Kim, Maarten Sap, Soham Dan, Thomas HartvigsenMELT @ COLM 2025 🇨🇦 PDF
  • PolyGuard: A Multilingual Safety Moderation Tool for 17 LanguagesPriyanshu Kumar, Devansh Jain, Akhila Yerukola, Liwei Jiang, Himanshu Beniwal, Thomas Hartvigsen, Maarten SapCOLM 2025 🇨🇦 PDF
  • COMMENTATOR: A Code-mixed Multilingual Text Annotation FrameworkRajvee Sheth, Shubh Nisar, Heenaben Prajapati, Himanshu Beniwal, Mayank SinghEMNLP 2024 · Demo 🇺🇸Core Rank: A* PDF Website 🕸️
  • PythonSaga: Redefining the Benchmark to Evaluate Code Generating LLMsAnkit Yadav, Himanshu Beniwal, Mayank SinghEMNLP 2024 🇺🇸Core Rank: A* PDF
  • Remember This Event That Year? 🤔 Assessing Temporal Information and Reasoning in Large Language ModelsHimanshu Beniwal, Dishant Patel, Kowsik Nandagopan D, Hritik Ladia, Ankit Yadav, Mayank SinghEMNLP 2024 🇺🇸Core Rank: A* PDF Website 🤔
  • Cross-lingual Editing in Multilingual Language ModelsHimanshu Beniwal, Kowsik Nandagopan D, Mayank SinghEACL 2024 🇲🇹Core Rank: A PDF Website 🕸️
  • Explainable Transformer-based Anomaly Detection for IoT SecurityAamir Saghir, Himanshu Beniwal, Kim Duc Tran, Ali Raza, Ludovic Koehl, Xianyi Zeng, Kim Phuc TranEAI SaSeIoT 2023International Conference on Safety and Security in IoT, pp. 83–109 · Springer Nature Switzerland Springer
  • A survey on near-human conversational agentsSatwinder Singh, Himanshu BeniwalJKSU-CIS 2021Journal of King Saud University — Computer and Information Sciences, 1319-1578, 2021 · IF: 13.473 (2021) PDF
  • Handwritten Digit Recognition using Machine LearningNarender Kumar, Himanshu BeniwalIJCSE 2018International Journal of Computer Sciences and Engineering, Vol. 06, Issue 05, pp. 96-100, 2018 · IF: 3.218 (2018) PDF Archived

Posters & Talks 🔊

  • [April 2025] Talk on “Cross-lingual Backdoors” at Plutous! [Recording] ⭐️
  • [March 2025] Talk on “GenAI in HealthCare”, at Google Developer Group - Silver Oak University, Ahmedabad, India 🇮🇳!!
  • [Sept 2024] Talk on “Editing Large Language Models”, MilaNLP, Italy! 🇮🇹
  • [March 2024] Talk on “Editing Large Language Models”, Google Research India, Bangalore, India.
  • [March 2024] PMRF Symposium 2024, “Cross-lingual Editing in Multilingual Language Models”, at IIT Indore.
  • [Feb 2024] [Research Week with Google 2024] at Google Research India, Bangalore, India, on ‘XME: Cross-lingual Model Editing in LLMs’.
  • [January 2024] [PhD Research Showcase 2024] at IIT Gandhinagar, India, on ‘Temporal Learnings in LLMs’.
  • [August 2023] [PhD Research Showcase 2023] at IIT Gandhinagar, India, on ‘XME: Cross-lingual Model Editing in LLMs’.
  • [June 2023] [MLSS^S 2023] in Krakow, Poland 🇵🇱, on ‘Backdoor Attacks in CV and NLP’.
  • [Feb 2023] Talk on “Backdoor Attacks in NLP”, at IISER Bhopal.

See the poster & talk gallery 🤩

Community Experience 👷🏻‍♂️

Teaching Assistantships 🛳️

Tula’s Institute January 2026 – May 2026 Dehradun, India 🇮🇳
  • Artificial Intelligence and Natural Language Processing, January 2026 to May 2026.
Dhirubhai Ambani Institute of Information and Communication Technology (DA-IICT) January 2023 – May 2025 Gandhinagar, India 🇮🇳
  • IT549: Deep Learning, 45+ students, Jan 2025 to May 2025, with Prof. Arpit Rana.
  • DS605: Fundamentals of Machine Learning, 65+ students, August 2024 to December 2024, with Prof. Arpit Rana.
  • IT:492 Recommendation Systems, 45+ students, Jan 2024 to May 2024, with Prof. Arpit Rana.
  • IT:496 Introduction to Data Mining, July 2023 to December 2023, with Prof. Arpit Rana.
  • IT:492 Recommendation Systems, Jan to May 2023, with Prof. Arpit Rana.
Indian Institute of Technology Gandhinagar August 2021 – December 2025 Gandhinagar, India 🇮🇳

Media Coverage 📰

Students Mentored 🧑🏻‍💻

Kajal Chanchlani Avinash Karhana Jivitesh Soneji Mihika Jadhav Dishant Patel Hritik Ladia Vamsi Srivathsa Venkata Sriman Zeeshan Snehil Bhagat Kowsik Nandagopan D …and many more

Notebooks 📒

Recommendation Systems

Introduction to Data Mining

Past Projects 👨🏻‍💻

Backdoor Attacks in Computer Vision Tasks Himanshu Beniwal, Prof. Shanmuganathan Raman August 2022 – December 2022 99% ASR @ 0.1% budget

Explored backdoor attack in MNIST, CIFAR10, MOT, and real-world datasets. Reporting, 99% attack success rate with 0.1% poisoning budget. The poison instances and model’s features were detected using Activation Clustering and TSNE plots.

Results

Detected people in frames captured from a real-world video
A. Captured frames from the real-world video.
Detected people in frames from the MOT17 dataset
B. Captured frames from the MOT dataset.
Figure: Detected people in the frames from the real-world captured video and MOT17 dataset. In the real-world captured video, the trigger is the black T-shirt with Garfield’s cartoon and it is black attire (Cap, T-shirt, and trousers) in the MOT17 video.

Poisoning Attacks in Text Classification and Generation Himanshu Beniwal, Prof. Mayank Singh January 2023 – May 2023 99% ASR · 95% clean acc.

Experimented with clean-label and label-flipping attacks in text generation and classification. Achieving 99% ASR with 95% clean-accuracy on SST-2 for classification. Classification models with triggers: Google’, ‘James Bond’, and ‘cf’. Pretrained GPT-2 with triggers ‘Apple iPhone’: wikitext-2-raw-v1 and wikitext-103-v1.

Predictions from bert-base-uncased with and without the trigger word
Figure: Prediction from bert-base-uncased, without and with trigger ('Google'). The metrics were accuracy (95.60) and Attack Success Rate (99.63). Hosted on 🤗: himanshubeniwal/bert_cl_g_1700.

Assessing Empathetic Capabilities in Conversational Approaches Himanshu Beniwal, Prof. Satwinder Singh January 2021 – May 2021 M.Tech. Thesis

To assess the empathetic capabilities in conversational approaches using seq2seq and transformers variations like generative, bi-encoder, poly-encoder, and ranker for empathetic dialogue dataset.

Gold empathetic conversations generated by different architectures
Gold empathetic conversations from different architectures.

Last updated: August 24, 2026  ·  Photo gallery 📸