Evaluation, Reliability & Deployment of Language and Vision-Language Models
2024 — Present
Ph.D. research, Rensselaer Polytechnic Institute
I develop methods for measuring AI reliability beyond task accuracy. In Unified Deployment-Aware Evaluation of Open Reasoning Language Models (TMLR, 2026), I evaluated seven model configurations across four benchmarks and three prompting strategies, covering 84 conditions and 19,992 examples, while jointly analyzing performance, uncertainty, latency, memory, prompt sensitivity, and deployment trade-offs.
I also developed HypothesisMed for structured reliability reporting, MILU for multimodal semantic evaluation, and earlier comparative work on ChatGPT versus DeepSeek for AI-based code generation.
Provenance, Auditability & Secure Knowledge Systems
2020 — Present
M.Sc. foundation at KUET; continuing Ph.D. research at RPI
I develop provenance and integrity mechanisms that make digital and AI-generated knowledge traceable, tamper-evident, reproducible, and securely exchanged. SlideChain provides blockchain-backed semantic provenance for structured VLM outputs over 1,117 lecture slides, with 100% tamper detection and deterministic reproducibility under controlled testing.
Earlier work includes my M.Sc. research on secure multi-party skyline queries and ShaEr, a privacy-preserving medical-data sharing framework.
Safe Agentic AI & Scientific Knowledge Governance
2025 — Present
Ph.D. research, Rensselaer Polytechnic Institute
I study how autonomous AI systems behave and act under permissions, policies, social interaction, and governance constraints. My workflow-agent research evaluates permission scope, semantic policy gating, and tamper-evident execution records; OpenClaw studies risky instruction propagation and norm enforcement across 14,490 autonomous agents; and ADAPT investigates adaptive governance in AI-mediated scholarly publishing.
The ADAPT work led to U.S. Provisional Patent Application No. 63/975,609.
Multimodal AI for Education & Medical Workflows
2025 — Present
Ph.D. research and system development, Rensselaer Polytechnic Institute
I develop multimodal and interactive AI systems for education and medical knowledge workflows. I helped build MEDI-SLATE, comprising 1,117 slides and 262,182 narration tokens from a 23-lecture medical imaging course; contributed to ALIVE, a fully local interactive lecture system; and co-developed an optically emulated computed tomography scanner for hands-on medical imaging education.
Foundations in Machine Learning & Natural Language Processing
2019 — 2025
Earlier research, Khulna University of Engineering & Technology and RPI
My earlier work established a foundation in machine learning, NLP, and robustness, including Bangla and phonetic-Bangla sentiment analysis, unified sentiment and emotion recognition, adversarially robust text classification, and closed-domain question answering. Related methodological work includes N-ReLU, a stochastic extension of the ReLU activation function.