Publications - Current Year

2026

  1. Conference paper
    D2
    “Images as Tables: In-Context Learning with TabPFN for Low-Data Detection of AI-Generated Images,” in 2nd ICML Workshop on Foundation Models for Structured Data (FMSD 2026), Seoul, Korea, 2026.
  2. Conference paper
    D2
    “MixFlow: Mixed Source Distributions Improve Rectified Flows,” in 2nd Workshop on Deep Generative Model in Machine Learning: Theory, Principle and Efficacy (ICLR 2026 DeLTa Workshop), Rio de Janeiro, Brazil, 2026.
  3. Conference paper
    D2
    “CFM: Language-aligned Concept Foundation Model for Vision,” in Computer Vision -- ECCV 2026, Malmö, Sweden.
  4. Article
    D2
    “STELLA: a modular framework for SpatioTemporal Event-based Lagrangian particLe trAcking,” Experiments in Fluids, vol. 67, 2026.
  5. Conference paper
    D2
    “TEVI: Text-Conditioned Editing of Visual Representations via Sparse Autoencoders for Improved Vision-Language Alignment,” in Findings of EMNLP 2026, Budapest, Hungary.
  6. Conference paper
    D2
    “More Images, More Problems? A Controlled Analysis of VLM Failure Modes,” in Findings of the Association for Computational Linguistics (ACL 2026), San Diego, CA, USA, 2026.
  7. Conference paper
    D2
    “Position: We need to re-think the concept of ‘real’ images,” in Forty-third International Conference on Machine Learning Position Paper Track (ICML 2026), Seoul, Korea, 2026.
  8. Conference paper
    D2
    “Align Once to Explain: Feature Alignment for Scalable B-cosification of Foundational Vision Transformers,” in IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR 2026), Denver, CO, USA, 2026.
  9. Conference paper
    D2
    “MM-TS: Multi-Modal Temperature and Margin Schedules for Contrastive Learning with Long-Tail Data,” in IEEE/CVF Winter Conference on Applications of Computer Vision (WACV 2026), Tucson, AZ, USA, 2026.
  10. Article
    D2
    “Sports-QA: A Large-Scale Video Question Answering Benchmark for Complex and Professional Sports,” International Journal of Computer Vision, vol. 134, 2026.
  11. Conference paper
    D2
    “GeoDiv: Framework for Measuring Geographical Diversity in Text-to-Image Models,” in The Fourteenth International Conference on Learning Representations (ICLR 2026), Rio de Janeiro, Brazil, 2026.
  12. Article
    D2
    “CROC: Evaluating and Training T2I Metrics with Pseudo- and Human-Labeled Contrastive Robustness Checks,” Transactions of the Association for Computational Linguistics, vol. 14, 2026.
  13. Article
    D2
    “Parameterized Adverse Lens Corruptions to Probe Model Robustness to Optical Tolerances,” Transactions on Machine Learning Research, vol. 2026, no. 7, 2026.
  14. Article
    D2
    “Beyond Accuracy: What Matters in Designing Well-Behaved Models?,” Transactions on Machine Learning Research, vol. 2026, no. 1, 2026.
  15. Article
    D2
    “What’s in the Bottle? A Survey and Roadmap of Concept Bottleneck Models,” Transactions on Machine Learning Research, vol. 2026, no. 5, 2026.
  16. Paper
    D2
    “Certified Circuits: Stability Guarantees for Mechanistic Circuits,” 2026. [Online]. Available: https://arxiv.org/abs/2602.22968.
  17. Paper
    D2
    “SceneTok: A Compressed, Diffusable Token Space for 3D Scenes,” 2026. [Online]. Available: https://arxiv.org/abs/2602.18882.
  18. Paper
    D2
    “Interpretability Without Tradeoffs: Disentangling Polysemanticity At Equal Predictive Performance,” 2026. [Online]. Available: https://arxiv.org/abs/2605.31304.
  19. Paper
    D2
    “SemanticNVS: Improving Semantic Scene Understanding in Generative Novel View Synthesis,” 2026. [Online]. Available: https://arxiv.org/abs/2602.20079.
  20. Paper
    D2
    “Do Instance Priors Help Weakly Supervised Semantic Segmentation?,” 2026. [Online]. Available: https://arxiv.org/abs/2604.11170.
  21. Paper
    D6D2
    “SOCO: Benchmarking Semantic Object Correspondence in Vision Foundation Models,” 2026. [Online]. Available: https://arxiv.org/abs/2605.31597.
  22. Paper
    D2
    “Rewis3d: Reconstruction Improves Weakly-Supervised Semantic Segmentation,” 2026. [Online]. Available: https://arxiv.org/abs/2603.06374.
  23. Paper
    D2
    “DataComp-VLM: Improved Open Datasets for Vision-Language Models,” 2026. [Online]. Available: https://arxiv.org/abs/2606.28551.
  24. Paper
    D2
    “RAWDet-7: A Multi-Scenario Benchmark for Object Detection and Description on Quantized RAW Images,” 2026. [Online]. Available: https://arxiv.org/abs/2602.03760.
  25. Paper
    D2
    “CSFlow: Aligning Flow Matching with Human Contrast Sensitivity,” 2026. [Online]. Available: https://arxiv.org/abs/2606.08833.
  26. Paper
    D2
    “What is Missing? Explaining Neurons Activated by Absent Concepts,” 2026. [Online]. Available: https://arxiv.org/abs/2603.09787.
  27. Paper
    D2
    “What Matters for Scalable and Robust Learning in End-to-End Driving Planners?,” 2026. [Online]. Available: https://arxiv.org/abs/2603.15185.
  28. Paper
    D2
    “Predictive Query Language: A Domain-Specific Language for Predictive Modeling on Relational Databases,” 2026. [Online]. Available: https://arxiv.org/abs/2602.09572.
  29. Paper
    D2
    “PARCEL: Pool-Anchored Resampling with Conditioned Elastic Queries for Efficient Vision-Language Understanding,” 2026. [Online]. Available: https://arxiv.org/abs/2605.30126.
  30. Paper
    D2
    “ClipTTT: CLIP-Guided Test-Time Training Helps LVLMs See Better,” 2026. [Online]. Available: https://arxiv.org/abs/2603.26486.
  31. Paper
    D2
    “ActionPlan: Future-Aware Streaming Motion Synthesis via Frame-Level Action Planning,” 2026. [Online]. Available: https://arxiv.org/abs/2603.13500.
  32. Paper
    D2
    “From Codebooks to VLMs: Evaluating Automated Visual Discourse Analysis for Climate Change on Social Media,” 2026. [Online]. Available: https://arxiv.org/abs/2604.21786.
  33. Paper
    D2D6
    “EmoteGPT: 3D Human Facial Expressions from Natural Language Descriptions,” 2026. [Online]. Available: https://arxiv.org/abs/2607.02674.
  34. Paper
    D2
    “DAVE: Distribution-aware Attribution via ViT Gradient Decomposition,” 2026. [Online]. Available: https://arxiv.org/abs/2602.06613.
  35. Paper
    D2
    “SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models,” 2026. [Online]. Available: https://arxiv.org/abs/2604.20705.
  36. Paper
    D2
    “R-CoV: Region-Aware Chain-of-Verification for Alleviating Object Hallucinations in LVLMs,” 2026. [Online]. Available: https://arxiv.org/abs/2604.20696.
  37. Paper
    D2
    “ClimateVID -- Social Media Videos Analysis and Challenges Involved,” 2026. .
  38. Paper
    D2
    “Seeing Through Circuits: Faithful Mechanistic Interpretability for Vision Transformers,” 2026. [Online]. Available: https://arxiv.org/abs/2604.14477.

2025

  1. Conference paper
    D2
    “y-Quant: Towards Learnable Quantization for Low-bit Pattern Recognition,” in Pattern Recognition (DAGM GCPR 2025), Freiburg, Germany, 2026.
  2. Conference paper
    D2
    “MT-Occ: Single-View 3D Occupancy Prediction via Multi-task Distillation,” in Pattern Recognition (DAGM GCPR 2025), Freiburg, Germany, 2026.