SWAN AI Research Group and collaborators at USC Viterbi

about

Ph.D. Candidate · Georgia State University · Atlanta

Making Multimodal Agents and World Action Models reliable.

I am a fifth-year Ph.D. candidate in Computer Science at Georgia State University, USA, advised by Prof. Ugur Kursuncu in the SWAN AI Research Group. During my Ph.D., I spent the summers of 2025 and 2026 as an Applied Scientist Intern on Siemens’ Data & AI Research team in Seattle, working on multimodal industrial foundation models and synthetic data generation and evaluation for Physical AI. In summer 2024, I interned with the Neuro-Symbolic Computing and Intelligence group at SRI International, researching uncertainty quantification for multimodal LLMs through an ARPA-H-funded project.

My research focuses on building reliable multimodal AI models and agents that learn at test time, ground their reasoning in evidence, improve from their own experience, and remain safe as tasks and environments change. This work has appeared at ACL, AAAI ICWSM, IEEE BigData and ACM TIST.

  • Grounded Multimodal Reasoning: Combining structured knowledge, visual grounding, and knowledge distillation to improve contextual understanding in vision-language models (ACL Findings ‘25, IEEE BigData ‘24).
  • Agent Safety: Investigating vulnerabilities through multi-turn red-teaming, adversarial interactions, and theory-guided behavioral evaluation (AAAI ICWSM ‘27, accepted, ACM TIST ‘26).
  • Uncertainty and Interpretability: Calibrating multimodal confidence using visual evidence (EMNLP GroundLM Workshop ‘26) and combining conformal prediction with representation probing for early failure detection in agents (ICML Workshop ‘26).
  • Self-Improving Agents: Learning from failures through preference optimization and co-evolving agents (Preprint).

I am also exploring world-action models for Physical AI, multi-agent simulation and alignment with human behavior, and enterprise agents that help users discover, access, and reason over organizational data.

research interests

  • World Action Models
  • Test-Time Learning
  • Neuro-Symbolic Systems
  • Knowledge Graphs
  • Knowledge Acquisition
  • Physical AI
  • Interpretability
  • Uncertainty Quantification
  • Automated Hypothesis Extraction

Collaborators and present & past affiliations

Siemens SRI International Amazon Adobe Meta Salesforce Rakuten Bosch NVIDIA Georgia Tech Wright State University Cal Poly San Luis Obispo

news

Aug 31, 2026 Cross-Modal Grounding for Calibrated Confidence in Vision Language Models has been accepted at EMNLP 2026 GroundLM Workshop (Conference on Empirical Methods in Natural Language Processing). See you in Budapest, Hungary :tada:
Aug 31, 2026 Attending ACM AI Leadership Summit 2026 in Atlanta, Georgia from August 31 to September 2, 2026. Looking forward to connecting with fellow researchers and industry leaders in AI! :tada:
Aug 30, 2026 Echoes of Human Malice in Agents: Benchmarking LLMs for Multi-Turn Online Harassment Attacks has been accepted at ICWSM 2027 (AAAI International Conference on Web and Social Media). See you in Edinburgh, Scotland :tada:
Jun 15, 2026 From Actions to Understanding: Conformal Interpretability of Temporal Concepts in LLM Agents has been accepted at the Mechanistic Interpretability Workshop, ICML 2026 in Seoul, South Korea. :tada:
May 26, 2026 Excited to share that our tutorial, “Knowledge-Infused Multimodal Learning,” has been accepted at ICWSM ’26, the 20th International AAAI Conference on Web and Social Media! I will be presenting at the University of Southern California, USC Information Sciences Institute, Los Angeles, alongside Agnik Saha and Professor Ugur Kursuncu. We will discuss some of our lab’s recent work on knowledge graph construction and how to design vision-language, knowledge-guided frameworks for more reliable and interpretable multimodal AI systems. Tutorial website
May 15, 2026 Excited to be spending the summer in Seattle, Washington, for my second stint at Siemens Data & AI Research. I’m looking forward to collaborating on Physical AI initiatives to help design and develop the next generation of AI systems. If you’re in the Seattle area, please feel free to reach out :tada:
Apr 01, 2026 New preprint: From Actions to Understanding: Conformal Interpretability of Temporal Concepts in LLM Agents, on step-wise conformal probes for early failure detection in LLM agents. :chart_with_upwards_trend:
Feb 01, 2026 Playing Devil’s Advocate: Unmasking Toxicity and Vulnerabilities in Large Vision-Language Models has been accepted at ACM Transactions on Intelligent Systems and Technology (TIST).
Nov 27, 2025 New preprint with collaborators at KAIST and Amazon: Co-Evolving Agents: Learning from Failures as Hard Negatives.
Oct 16, 2025 New preprint: Echoes of Human Malice in Agents, a benchmark for multi-turn online harassment attacks on LLM agents.
May 15, 2025 Started as an Applied Scientist Intern at Siemens in Seattle, on the Data & AI Research team, working on multimodal industrial foundation models.
May 15, 2025 Just KIDDIN’: Knowledge Infusion and Distillation for Detection of INdecent Memes has been accepted to ACL 2025 Findings (acceptance rate 19.1%). Grateful for the ACL Student Travel Award. :tada:
Dec 15, 2024 Enhancing Cross-Modal Contextual Congruence for Crowdfunding Success was presented at IEEE BigData 2024 (acceptance rate 18.7%), supported by the IEEE BigData Student Travel Award.

I earned my Bachelor’s degree in Electronics and Communication Engineering from Veer Surendra Sai University of Technology in 2020. Before starting my Ph.D., I was a Data Scientist in Rakuten’s Data Science Consulting team, working on business impact estimation, customer acquisition, and funnel analysis across the Rakuten group. I collaborated with an international team and developed interpretable churn prediction models to inform customer retention strategies. I also held research internships at Bosch Research and NVIDIA Research.

education

PhD in Computer Science
2022 – present
B.Tech. in Electronics and Communication Engineering
Academic Scholarship
2016 – 2020

experience

Applied Scientist Intern, Data & AI Research
Siemens Seattle, WA
Synthetic data generation and evaluation for Physical AI foundation models; pretraining and post-training of multimodal industrial foundation models
Summers 2025 & 2026
Research Intern, Neuro-Symbolic Computing and Intelligence
SRI International Menlo Park, CA
Uncertainty quantification for multimodal LLMs using visual grounding, on an ARPA-H funded project
Summer 2024
Data Scientist
Rakuten Bengaluru, India
Churn prediction and customer acquisition modelling for the Data Science Consulting group
2021 – 2022
Research Intern
Bosch Research Bengaluru, India
Deep learning research
2020
Research Intern
NVIDIA Bengaluru, India
GPU-accelerated computing and deep learning
2018 – 2019

selected publications

  1. ICML WS
    conformal.png
    From Actions to Understanding: Conformal Interpretability of Temporal Concepts in LLM Agents
    Trilok Padhi, Ramneet Kaur, Krishiv Agarwal, and 9 more authors
    In Mechanistic Interpretability Workshop, 43rd International Conference on Machine Learning (ICML), Seoul, South Korea, 2026
  2. ACM TIST
    devils.png
    Playing Devil’s Advocate: Unmasking Toxicity and Vulnerabilities in Large Vision-Language Models
    Abdulkadir Erol, Trilok Padhi, Agnik Saha, and 2 more authors
    ACM Transactions on Intelligent Systems and Technology, 2026
  3. Preprint
    coevolving.png
    Co-Evolving Agents: Learning from Failures as Hard Negatives
    Yeonsung Jung, Trilok Padhi, Sina Shaham, and 4 more authors
    arXiv preprint arXiv:2511.22254 (under review), 2025
  4. ICWSM
    echoes.png
    Echoes of Human Malice in Agents: Benchmarking LLMs for Multi-Turn Online Harassment Attacks
    Trilok Padhi, Pinxian Lu, Abdulkadir Erol, and 5 more authors
    In Proceedings of the International AAAI Conference on Web and Social Media (ICWSM) (equal contribution with P. Lu), 2027
  5. ACL Findings
    kiddin.png
    Just KIDDIN’: Knowledge Infusion and Distillation for Detection of INdecent Memes
    Rahul Garg, Trilok Padhi, Hemang Jain, and 2 more authors
    In Findings of the Association for Computational Linguistics: ACL 2025 (equal contribution with R. Garg; acceptance rate 19.1%), 2025
  6. EMNLP WS
    seeing.png
    Cross-Modal Grounding for Calibrated Confidence in Vision Language Models
    Trilok Padhi, Ramneet Kaur, Adam D. Cobb, and 7 more authors
    In GroundLM Workshop, Conference on Empirical Methods in Natural Language Processing (EMNLP), Budapest, Hungary (the arXiv preprint appears under its earlier title), 2026
  7. IEEE BigData
    congruence.png
    Enhancing Cross-Modal Contextual Congruence for Crowdfunding Success using Knowledge-infused Learning
    Trilok Padhi, Ugur Kursuncu, Yaman Kumar, and 2 more authors
    In 2024 IEEE International Conference on Big Data (BigData) (acceptance rate 18.7%), 2024