Ph.D. Candidate · Georgia State University · Atlanta
Making Multimodal Agents and World Action Models reliable.
I am a fifth-year Ph.D. candidate in Computer Science at Georgia State University, USA, advised by Prof. Ugur Kursuncu in the SWAN AI Research Group. During my Ph.D., I spent the summers of 2025 and 2026 as an Applied Scientist Intern on Siemens’ Data & AI Research team in Seattle, working on multimodal industrial foundation models and synthetic data generation and evaluation for Physical AI. In summer 2024, I interned with the Neuro-Symbolic Computing and Intelligence group at SRI International, researching uncertainty quantification for multimodal LLMs through an ARPA-H-funded project.
My research focuses on building reliable multimodal AI models and agents that learn at test time, ground their reasoning in evidence, improve from their own experience, and remain safe as tasks and environments change. This work has appeared at ACL, AAAI ICWSM, IEEE BigData and ACM TIST.
Grounded Multimodal Reasoning: Combining structured knowledge, visual grounding, and knowledge distillation to improve contextual understanding in vision-language models (ACL Findings ‘25, IEEE BigData ‘24).
Agent Safety: Investigating vulnerabilities through multi-turn red-teaming, adversarial interactions, and theory-guided behavioral evaluation (AAAI ICWSM ‘27, accepted, ACM TIST ‘26).
Uncertainty and Interpretability: Calibrating multimodal confidence using visual evidence (EMNLP GroundLM Workshop ‘26) and combining conformal prediction with representation probing for early failure detection in agents (ICML Workshop ‘26).
Self-Improving Agents: Learning from failures through preference optimization and co-evolving agents (Preprint).
I am also exploring world-action models for Physical AI, multi-agent simulation and alignment with human behavior, and enterprise agents that help users discover, access, and reason over organizational data.
Cross-Modal Grounding for Calibrated Confidence in Vision Language Models has been accepted at EMNLP 2026 GroundLM Workshop (Conference on Empirical Methods in Natural Language Processing). See you in Budapest, Hungary
Aug 31, 2026
Attending ACM AI Leadership Summit 2026 in Atlanta, Georgia from August 31 to September 2, 2026. Looking forward to connecting with fellow researchers and industry leaders in AI!
Aug 30, 2026
Echoes of Human Malice in Agents: Benchmarking LLMs for Multi-Turn Online Harassment Attacks has been accepted at ICWSM 2027 (AAAI International Conference on Web and Social Media). See you in Edinburgh, Scotland
Jun 15, 2026
From Actions to Understanding: Conformal Interpretability of Temporal Concepts in LLM Agents has been accepted at the Mechanistic Interpretability Workshop, ICML 2026 in Seoul, South Korea.
May 26, 2026
Excited to share that our tutorial, “Knowledge-Infused Multimodal Learning,” has been accepted at ICWSM ’26, the 20th International AAAI Conference on Web and Social Media! I will be presenting at the University of Southern California, USC Information Sciences Institute, Los Angeles, alongside Agnik Saha and Professor Ugur Kursuncu. We will discuss some of our lab’s recent work on knowledge graph construction and how to design vision-language, knowledge-guided frameworks for more reliable and interpretable multimodal AI systems. Tutorial website
May 15, 2026
Excited to be spending the summer in Seattle, Washington, for my second stint at Siemens Data & AI Research. I’m looking forward to collaborating on Physical AI initiatives to help design and develop the next generation of AI systems. If you’re in the Seattle area, please feel free to reach out
Playing Devil’s Advocate: Unmasking Toxicity and Vulnerabilities in Large Vision-Language Models has been accepted at ACM Transactions on Intelligent Systems and Technology (TIST).
Started as an Applied Scientist Intern at Siemens in Seattle, on the Data & AI Research team, working on multimodal industrial foundation models.
May 15, 2025
Just KIDDIN’: Knowledge Infusion and Distillation for Detection of INdecent Memes has been accepted to ACL 2025 Findings (acceptance rate 19.1%). Grateful for the ACL Student Travel Award.
Dec 15, 2024
Enhancing Cross-Modal Contextual Congruence for Crowdfunding Success was presented at IEEE BigData 2024 (acceptance rate 18.7%), supported by the IEEE BigData Student Travel Award.
I earned my Bachelor’s degree in Electronics and Communication Engineering from Veer Surendra Sai University of Technology in 2020. Before starting my Ph.D., I was a Data Scientist in Rakuten’s Data Science Consulting team, working on business impact estimation, customer acquisition, and funnel analysis across the Rakuten group. I collaborated with an international team and developed interpretable churn prediction models to inform customer retention strategies. I also held research internships at Bosch Research and NVIDIA Research.
Large Language Models (LLMs) are increasingly deployed as autonomous agents capable of reasoning, planning, and acting within interactive environments. Despite their growing capability to perform multi-step reasoning and decision-making tasks, internal mechanisms guiding their sequential behavior remain opaque. This paper presents a framework for interpreting the temporal evolution of concepts in LLM agents through a step-wise conformal lens. We introduce the conformal interpretability framework for temporal tasks, which combines step-wise reward modeling with conformal prediction to statistically label model’s internal representation at each step as successful or failing. Linear probes are then trained on these representations to identify directions of temporal concepts - latent directions in the model’s activation space that correspond to consistent notions of success, failure or reasoning drift. Experimental results on two simulated interactive environments, namely ScienceWorld and AlfWorld, demonstrate that these temporal concepts are linearly separable, revealing interpretable structures aligned with task success. We further show preliminary results on improving an LLM agent’s performance by leveraging the proposed framework for steering the identified successful directions inside the model. The proposed approach, thus, offers a principled method for early failure detection as well as intervention in LLM-based agents, paving the path towards trustworthy autonomous language models in complex interactive settings.
@inproceedings{padhi2026actions,title={From Actions to Understanding: Conformal Interpretability of Temporal Concepts in LLM Agents},author={Padhi, Trilok and Kaur, Ramneet and Agarwal, Krishiv and Cobb, Adam D. and Elenius, Daniel and Acharya, Manoj and Samplawski, Colin and Berenbeim, Alexander M. and Bastian, Nathaniel D. and Jha, Susmit and Kursuncu, Ugur and Roy, Anirban},booktitle={Mechanistic Interpretability Workshop, 43rd International Conference on Machine Learning (ICML)},location={Seoul, South Korea},year={2026},}
ACM TIST
Playing Devil’s Advocate: Unmasking Toxicity and Vulnerabilities in Large Vision-Language Models
Abdulkadir Erol, Trilok Padhi, Agnik Saha, and 2 more authors
ACM Transactions on Intelligent Systems and Technology, 2026
The rapid advancement of Large Vision-Language Models (LVLMs) has enhanced capabilities offering potential applications from content creation to productivity enhancement. Despite their innovative potential, LVLMs exhibit vulnerabilities, especially in generating potentially toxic or unsafe responses. Malicious actors can exploit these vulnerabilities to propagate toxic content in an automated (or semi-) manner, leveraging the susceptibility of LVLMs to deception via strategically crafted prompts without fine-tuning or compute-intensive procedures. Despite the red-teaming efforts and inherent potential risks associated with the LVLMs, exploring vulnerabilities of LVLMs remains nascent and yet to be fully addressed in a systematic manner. This study systematically examines the vulnerabilities of open-source LVLMs, including LLaVA, InstructBLIP, Fuyu, and Qwen, using adversarial prompt strategies that simulate real-world social manipulation tactics informed by social theories. Our findings show that (i) toxicity and insulting are the most prevalent behaviors, with the mean rates of 16.13% and 9.75%, respectively; (ii) Qwen-VL-Chat, LLaVA-v1.6-Vicuna-7b, and InstructBLIP-Vicuna-7b are the most vulnerable models, exhibiting toxic response rates of 21.50%, 18.30% and 17.90%, and insulting responses of 13.40%, 11.70% and 10.10%, respectively; (iii) prompting strategies incorporating dark humor and multimodal toxic prompt completion significantly elevated these vulnerabilities. Despite being fine-tuned for safety, these models still generate content with varying degrees of toxicity when prompted with adversarial inputs, highlighting the urgent need for enhanced safety mechanisms and robust guardrails in LVLM development.
@article{erol2026devils,title={Playing Devil's Advocate: Unmasking Toxicity and Vulnerabilities in Large Vision-Language Models},author={Erol, Abdulkadir and Padhi, Trilok and Saha, Agnik and Kursuncu, Ugur and Aktas, Mehmet Emin},journal={ACM Transactions on Intelligent Systems and Technology},publisher={Association for Computing Machinery},year={2026},}
Preprint
Co-Evolving Agents: Learning from Failures as Hard Negatives
Yeonsung Jung, Trilok Padhi, Sina Shaham, and 4 more authors
The rapid progress of large foundation models has accelerated the development of task-specialized agents across diverse domains. However, the effectiveness of agents remains tightly coupled with the quality of training data, while curating task-specific datasets remains costly and often infeasible in real-world scenarios. Recent work has explored self-improving agents that autonomously generate, refine, and re-train on their own trajectories. A prominent line of approaches further leverages preference optimization by pairing predicted trajectories with scarce ground-truth trajectories, enabling agents to learn directly from their own failures. While these methods outperform supervised fine-tuning, their heavy reliance on predicted trajectories under limited ground-truth supervision leaves them prone to overfitting. To address this, we propose a co-evolving agents framework in which a target agent improves jointly with an auxiliary failure agent. The failure agent learns through preference optimization over failure trajectories from both the target and itself, thereby generating hard negatives that are close to success yet remain failures. Incorporating these informative hard negatives into the target agent’s optimization sharpens decision boundaries and enhances generalization. Our comprehensive analysis and experiments across benchmark datasets show that our method not only shows improved performance but also demonstrates that failures, instead of being used as-is, can be systematically transformed into structured and valuable learning signals in self-improving agents.
@article{jung2025coevolving,title={Co-Evolving Agents: Learning from Failures as Hard Negatives},author={Jung, Yeonsung and Padhi, Trilok and Shaham, Sina and Khullar, Dipika and Jeong, Joonhyun and Mehrabi, Ninareh and Yang, Eunho},journal={arXiv preprint arXiv:2511.22254},year={2025},}
ICWSM
Echoes of Human Malice in Agents: Benchmarking LLMs for Multi-Turn Online Harassment Attacks
Trilok Padhi, Pinxian Lu, Abdulkadir Erol, and 5 more authors
In Proceedings of the International AAAI Conference on Web and Social Media (ICWSM) (equal contribution with P. Lu), 2027
Large Language Model (LLM) agents are powering a growing share of interactive web applications, yet remain vulnerable to misuse and harm. Prior jailbreak research has largely focused on single-turn prompts, whereas real harassment often unfolds over multi-turn interactions. In this work, we present the Online Harassment Agentic Benchmark consisting of: (i) a synthetic multi-turn harassment conversation dataset, (ii) a multi-agent (e.g., harasser, victim) simulation informed by repeated game theory, (iii) three jailbreak methods attacking agents across memory, planning, and fine-tuning, and (iv) a mixed-methods evaluation framework. We utilize two prominent LLMs, LLaMA-3.1-8B-Instruct (open-source) and Gemini-2.0-flash (closed-source). Our results show that jailbreak tuning makes harassment nearly guaranteed with an attack success rate of 95.78–96.89% vs. 57.25–64.19% without tuning in Llama, and 99.33% vs. 98.46% without tuning in Gemini, while sharply reducing refusal rate to 1-2% in both models. The most prevalent toxic behaviors are Insult with 84.9–87.8% vs. 44.2–50.8% without tuning, and Flaming with 81.2–85.1% vs. 31.5–38.8% without tuning, indicating weaker guardrails compared to sensitive categories such as sexual or racial harassment. Qualitative evaluation further reveals that attacked agents reproduce human-like aggression profiles, such as Machiavellian/psychopathic patterns under planning, and narcissistic tendencies with memory. Counterintuitively, closed-source and open-source models exhibit distinct escalation trajectories across turns, with closed-source models showing significant vulnerability. Overall, our findings show that multi-turn and theory-grounded attacks not only succeed at high rates but also mimic human-like harassment dynamics, motivating the development of robust safety guardrails to ultimately keep online platforms safe and responsible.
@inproceedings{padhi2027echoes,title={Echoes of Human Malice in Agents: Benchmarking LLMs for Multi-Turn Online Harassment Attacks},author={Padhi, Trilok and Lu, Pinxian and Erol, Abdulkadir and Sutar, Tanmay and Sharma, Gauri and Sonmez, Mina and De Choudhury, Munmun and Kursuncu, Ugur},booktitle={Proceedings of the International AAAI Conference on Web and Social Media (ICWSM)},publisher={Association for the Advancement of Artificial Intelligence},year={2027},}
ACL Findings
Just KIDDIN’: Knowledge Infusion and Distillation for Detection of INdecent Memes
Rahul Garg, Trilok Padhi, Hemang Jain, and 2 more authors
In Findings of the Association for Computational Linguistics: ACL 2025 (equal contribution with R. Garg; acceptance rate 19.1%), 2025
Toxicity identification in online multimodal environments remains a challenging task due to the complexity of contextual connections across modalities (e.g., textual and visual). In this paper, we propose a novel framework that integrates Knowledge Distillation (KD) from Large Visual Language Models (LVLMs) and knowledge infusion to enhance the performance of toxicity detection in hateful memes. Our approach extracts sub-knowledge graphs from ConceptNet, a large-scale commonsense Knowledge Graph (KG) to be infused within a compact VLM framework. The relational context between toxic phrases in captions and memes, as well as visual concepts in memes enhance the model’s reasoning capabilities. Experimental results from our study on two hate speech benchmark datasets demonstrate superior performance over the state-of-the-art baselines across AU-ROC, F1, and Recall with improvements of 1.1%, 7%, and 35%, respectively. Given the contextual complexity of the toxicity detection task, our approach showcases the significance of learning from both explicit (i.e. KG) as well as implicit (i.e. LVLMs) contextual cues incorporated through a hybrid neurosymbolic approach. This is crucial for real-world applications where accurate and scalable recognition of toxic content is critical for creating safer online environments.
@inproceedings{garg2025kiddin,title={Just KIDDIN': Knowledge Infusion and Distillation for Detection of INdecent Memes},author={Garg, Rahul and Padhi, Trilok and Jain, Hemang and Kursuncu, Ugur and Kumaraguru, Ponnurangam},booktitle={Findings of the Association for Computational Linguistics: ACL 2025},publisher={Association for Computational Linguistics},year={2025},}
EMNLP WS
Cross-Modal Grounding for Calibrated Confidence in Vision Language Models
Trilok Padhi, Ramneet Kaur, Adam D. Cobb, and 7 more authors
In GroundLM Workshop, Conference on Empirical Methods in Natural Language Processing (EMNLP), Budapest, Hungary (the arXiv preprint appears under its earlier title), 2026
We introduce a novel approach for calibrating uncertainty quantification (UQ) tailored for multi-modal large language models (LLMs). Existing state-of-the-art UQ methods rely on consistency among multiple responses generated by the LLM on an input query under diverse settings. However, these approaches often report higher confidence in scenarios where the LLM is consistently incorrect. This leads to a poorly calibrated confidence with respect to accuracy. To address this, we leverage cross-modal consistency in addition to self-consistency to improve the calibration of the multi-modal models. Specifically, we ground the textual responses to the visual inputs. The confidence from the grounding model is used to calibrate the overall confidence. Given that using a grounding model adds its own uncertainty in the pipeline, we apply temperature scaling - a widely accepted parametric calibration technique - to calibrate the grounding model’s confidence in the accuracy of generated responses. We evaluate the proposed approach across multiple multi-modal tasks, such as medical question answering (Slake) and visual question answering (VQAv2), considering multi-modal models such as LLaVA-Med and LLaVA. The experiments demonstrate that the proposed framework achieves significantly improved calibration on both tasks.
@inproceedings{padhi2026seeing,title={Cross-Modal Grounding for Calibrated Confidence in Vision Language Models},author={Padhi, Trilok and Kaur, Ramneet and Cobb, Adam D. and Acharya, Manoj and Roy, Anirban and Samplawski, Colin and Matejek, Brian and Berenbeim, Alexander M. and Bastian, Nathaniel D. and Jha, Susmit},booktitle={GroundLM Workshop, Conference on Empirical Methods in Natural Language Processing (EMNLP)},location={Budapest, Hungary},year={2026},}
IEEE BigData
Enhancing Cross-Modal Contextual Congruence for Crowdfunding Success using Knowledge-infused Learning
Trilok Padhi, Ugur Kursuncu, Yaman Kumar, and 2 more authors
In 2024 IEEE International Conference on Big Data (BigData) (acceptance rate 18.7%), 2024
The digital landscape continually evolves with multimodality, enriching the online experience for users. Creators and marketers aim to weave subtle contextual cues from various modalities into congruent content to engage users with a harmonious message. This interplay of multimodal cues is often a crucial factor in attracting users’ attention. However, this richness of multimodality presents a challenge to computational modeling, as the semantic contextual cues spanning across modalities need to be unified to capture the true holistic meaning of the multimodal content. This contextual meaning is critical in attracting user engagement as it conveys the intended message of the brand or the organization. In this work, we incorporate external commonsense knowledge from knowledge graphs to enhance the representation of multimodal data using compact Visual Language Models (VLMs) and predict the success of multi-modal crowdfunding campaigns. Our results show that external knowledge commonsense bridges the semantic gap between text and image modalities, and the enhanced knowledge-infused representations improve the predictive performance of models for campaign success upon the baselines without knowledge. Our findings highlight the significance of contextual congruence in online multimodal content for engaging and successful crowdfunding campaigns.
@inproceedings{padhi2024congruence,title={Enhancing Cross-Modal Contextual Congruence for Crowdfunding Success using Knowledge-infused Learning},author={Padhi, Trilok and Kursuncu, Ugur and Kumar, Yaman and Shalin, Valerie L. and Fronczek, Lane Peterson},booktitle={2024 IEEE International Conference on Big Data (BigData)},publisher={IEEE},year={2024},}