Tinted Frames: Question Framing Blinds Vision-Language Models

Wan-Cyuan Fan, Jiayun Luo, Declan Kutscher, Leonid Sigal, Ritwik Gupta

VLMs are selectively blind. They decide how much to look at an image based on question framing (open-ended vs Yes/No vs MCQ), even when the same visual reasoning is required. We analyze this phenomenon and mitigate it in this paper.

Artwork for Tinted Frames: Question Framing Blinds Vision-Language Models
TCBench: A Benchmark for Tropical Cyclone Track and Intensity Forecasting at the Global Scale

Milton Gomez, Marie McGraw, Saranya Ganesh S., Frederick Iat-Hin Tam, Ilia Azizi, Samuel Darmon, Monika Feldmann, Stella Bourdin, Louis Poulain-Auzéau, Suzana J. Camargo, Jonathan Lin, Dan Chavas, Chia-Ying Lee, Ritwik Gupta, Andrea Jenney, Tom Beucler

TCBench is a benchmark for evaluating global, short to medium-range (1-5 days) forecasts of tropical cyclone track and intensity. It builds on the IBTrACS observational dataset and includes state-of-the-art dynamical and neural weather models. Designed for accessibility, TCBench helps AI practitioners tackle domain-relevant TC challenges.

arXiv preprint paper github
Artwork for TCBench: A Benchmark for Tropical Cyclone Track and Intensity Forecasting at the Global Scale
The LLM Mirage: Economic Interests and the Subversion of Weaponization Controls

Ritwik Gupta, Andrew W. Reddie

U.S. AI security policy is increasingly shaped by an LLM Mirage, the belief that national security risks scale in proportion to the compute used to train frontier language models. That premise fails in two ways: it miscalibrates strategy because adversaries can obtain weaponizable capabilities with task-specific systems, and it destabilizes regulation because compute thresholds are easily renegotiated as domestic priorities shift.

ACM Conference on Fairness, Accountability, and Transparency (FAccT) 2026 paper
Artwork for The LLM Mirage: Economic Interests and the Subversion of Weaponization Controls
Crowdsourcing the Frontier: Advancing Hybrid Physics-ML Climate Simulation via a $50,000 Kaggle Competition

Jerry Lin, Zeyuan Hu, Tom Beucler, Katherine Frields, Hannah Christensen, Walter Hannah, Helge Heuer, Peter Ukkonnen, Laura A. Mansfield, Tian Zheng, Liran Peng, Ritwik Gupta, Pierre Gentine, Yusef Al-Naher, Mingjiang Duan, Kyo Hattori, Weiliang Ji, Chunhan Li, Kippei Matsuda, Naoki Murakami, Shlomo Ron, Marec Serlin, Hongjian Song, Yuma Tanabe, Daisuke Yamamoto, Jianyao Zhou, Mike Pritchard

Online stability in the low-resolution, real-geography setting is reproducibly achievable across diverse architectures, and offline and online zonal mean biases are near-identical across architectures.

Journal of Advances in Modeling Earth Systems paper github
Artwork for Crowdsourcing the Frontier: Advancing Hybrid Physics-ML Climate Simulation via a $50,000 Kaggle Competition
REOrdering Patches Improves Vision Models

Declan Kutscher, David M. Chan, Yutong Bai, Trevor Darrell, Ritwik Gupta

We show that long sequence models are sensitive to the order of patches provided to them, affecting performance by as much as 13%. We propose REOrder, an information-theoretic and RL-based approach to learning an optimal patch ordering to improve performance.

Neural Information Processing Systems (NeurIPS) 2025 website paper github
Artwork for REOrdering Patches Improves Vision Models
Enough Coin Flips Can Make LLMs Act Bayesian

Ritwik Gupta, Rodolfo Corona, Jiaxin Ge, Eric Wang, Dan Klein, Trevor Darrell, David M. Chan

Can LLMs accurately represent probabilities? We find in this work, no! By using biased coin flips as a simple but powerful exemplar, we show that in-context learning can be used to induce semi-accurate probability simulation in LLMs.

Association for Computational Linguistics (ACL) 2025 website paper github
Artwork for Enough Coin Flips Can Make LLMs Act Bayesian
ClimSim-Online: A Large Multi-Scale Dataset and Framework for Hybrid Physics-ML Climate Emulation

Sungduk Yu, Zeyuan Hu, Akshay Subramaniam, Walter Hannah, Liran Peng, Jerry Lin, Mohamed Aziz Bhouri, Ritwik Gupta, Björn Lütjens, Justus C Will, Gunnar Behrens, Julius JM Busecke, Nora Loose, Charles I Stern, Tom Beucler, Bryce Harrop, Helge Heuer, Benjamin R Hillman, Andrea Jenney, Nana Liu, Alistair White, Tian Zheng, Zhiming Kuang, Fiaz Ahmed, Elizabeth Barnes, Noah D Brenowitz, Christopher Bretherton, Veronika Eyring, Savannah Ferretti, Nicholas Lutsko, Pierre Gentine, Stephan Mandt, J David Neelin, Rose Yu, Laure Zanna, Nathan M Urban, Janni Yuval, Ryan Abernathey, Pierre Baldi, Wayne Chuang, Yu Huang, Fernando Iglesias-Suarez, Sanket Jantre, Po-Lun Ma, Sara Shamekh, Guang Zhang, Michael Pritchard

We introduce a significant new contribution to ClimSim, which provides a cross-platform, containerized pipeline to integrate ML models into operational climate simulators for hybrid testing.

Journal of Machine Learning Research (JMLR) 26 (2025) 1-85 paper
Artwork for ClimSim-Online: A Large Multi-Scale Dataset and Framework for Hybrid Physics-ML Climate Emulation
Whack-a-Chip: The Futility of Hardware-Centric Export Controls

Ritwik Gupta, Leah Walker, Andrew W. Reddie

We give the first, public evidence as to how leading PRC AI labs are effectively circumventing U.S. semiconductor export controls through better software. We question the basis and efficacy of the current export control regime.

arXiv preprint paper
Artwork for Whack-a-Chip: The Futility of Hardware-Centric Export Controls
Open-Source Assessments of AI Capabilities: The Proliferation of AI Analysis Tools, Replicating Competitor Models, and the Zhousidun Dataset

Ritwik Gupta, Leah Walker, Eli Glickman, Raine Koizumi, Sarthak Bhatnagar, Andrew W. Reddie

China is training machine learning models to target American and Allied navel vessels, but how well do they work? In this paper, we train a state-of-the-art machine learning model on a leaked Chinese dataset that labels Aegis combat system components on military vessels. We propose a new methodology for open source assessment of adversary AI capabilities.

BRSL Tech Report paper github
Artwork for Open-Source Assessments of AI Capabilities: The Proliferation of AI Analysis Tools, Replicating Competitor Models, and the Zhousidun Dataset
xT: Nested Tokenization for Larger Context in Large Images

Ritwik Gupta*, Shufan Li*, Tyler Zhu*, Jitendra Malik, Trevor Darrell, Karttikeya Mangalam

xT is a framework which lets you model extremely large images (upwards of 30,000 x 30,000 pixels) end-to-end on contemporary GPUs. You get higher accuracy with fewer parameters and less memory used per region.

International Conference on Machine Learning (ICML) 2024 website paper github
Artwork for xT: Nested Tokenization for Larger Context in Large Images
Russian Nuclear ASAT Weapons: The Fallout

Sarthak Bhatnagar, Eli Glickman, Bethany Goldblum, Ritwik Gupta, Kaitlyn Lenkeit, Jane Darby Menton, Andrew Neciuk, Andrew Reddie, Vishwaa Sofat, Leah Walker

What is the state of the existing space governance regime amid concerns that Moscow is developing a nuclear-tipped anti-satellite weapon in orbit?

Lawfare paper
Artwork for Russian Nuclear ASAT Weapons: The Fallout
LAION and the Challenges of Preventing AI-Generated CSAM

Ritwik Gupta

I examined the challenges in preventing AI-generated Child Sexual Abuse Material (CSAM), such as within the widely-used LAION-5B dataset, emphasizing the need for updated legal and technological strategies to tackle the spread of such content by generative AI technologies.

Tech Policy Press paper
Artwork for LAION and the Challenges of Preventing AI-Generated CSAM
Accelerating the Evolution of AI Export Controls

Ritwik Gupta, Andrew W. Reddie

Current US AI hardware export controls are based on the best AI accelerator chip available at that time. This presents wide loopholes which allow adversarial nations to still maintain their capabilities. We propose an alternate way to set export control thresholds based on the analysis of specific ML workloads.

Tech Policy Press paper
Artwork for Accelerating the Evolution of AI Export Controls
Confidence-Building Measures for Artificial Intelligence

Sarah Shoker, Andrew Reddie, Sarah Barrington, Ruby Booth, Miles Brundage, Husanjot Chahal, Michael Depp, Bill Drexel, Ritwik Gupta, Marina Favaro, Jake Hecla, Alan Hickey, Margarita Konaev, Kirthi Kumar, Nathan Lambert, Andrew Lohn, Cullen O'Keefe, Nazneen Rajani, Michael Sellitto, Robert Trager, Leah Walker, Alexa Wehsener, Jessica Young

Workshop proceedings from the the Confidence-Building Measures for Artificial Intelligence workshop hosted by the Geopolitics Team at OpenAI and the Berkeley Risk and Security Lab at the University of California.

Workshop proceedings paper
Artwork for Confidence-Building Measures for Artificial Intelligence
ClimSim: An open large-scale dataset for training high-resolution physics emulators in hybrid multi-scale climate simulators

Sungduk Yu, Walter M. Hannah, Liran Peng, Mohamed Aziz Bhouri, Ritwik Gupta, Jerry Lin, Björn Lütjens, Justus C. Will, Tom Beucler, Bryce E. Harrop, Benjamin R. Hillman, Andrea M. Jenney, Savannah L. Ferretti, Nana Liu, Anima Anandkumar, Noah D. Brenowitz, Veronika Eyring, Pierre Gentine, Stephan Mandt, Jaideep Pathak, Carl Vondrick, Rose Yu, Laure Zanna, Ryan P. Abernathey, Fiaz Ahmed, David C. Bader, Pierre Baldi, Elizabeth A. Barnes, Gunnar Behrens, Christopher S. Bretherton, Julius J. M. Busecke, Peter M. Caldwell, Wayne Chuang, Yilun Han, Yu Huang, Fernando Iglesias-Suarez, Sanket Jantre, Karthik Kashinath, Marat Khairoutdinov, Thorsten Kurth, Nicholas J. Lutsko, Po-Lun Ma, Griffin Mooers, J. David Neelin, David A. Randall, Sara Shamekh, Akshay Subramaniam, Mark A. Taylor, Nathan M. Urban, Janni Yuval, Guang J. Zhang, Tian Zheng, Michael S. Pritchard

The largest-ever dataset designed for hybrid ML-physics research. It comprises multi-scale climate simulations, developed by a consortium of climate scientists and ML researchers.

Best Paper Award Neural Information Processing Systems (NeurIPS) 2023 website paper github
Artwork for ClimSim: An open large-scale dataset for training high-resolution physics emulators in hybrid multi-scale climate simulators
Orbital hypersonic delivery systems threaten strategic stability

Ritwik Gupta

We assess that China's development of a fractional orbital hypersonic delivery system, combining hypersonic glide vehicles with orbital bombardment, presents a concerning challenge to global stability, allowing for faster, undetectable delivery of large nuclear payloads and signaling renewed interest in first-strike capabilities.

The Bulletin of Atomic Scientists paper
Artwork for Orbital hypersonic delivery systems threaten strategic stability
Scale-MAE: A Scale-Aware Masked Autoencoder for Multiscale Geospatial Representation Learning

Ritwik Gupta*, Colorado Reed*, Shufan Li*, Sarah Brockman, Christopher Funk, Brian Clipp, Kurt Keutzer, Salvatore Candido, Matt Uyttendaele, Trevor Darrell

A pre-training method to make encoders robust to imagery captured at varying satellite resolutions. State-of-the-art multi-scale pre-training method and the largest satellite imagery foundation model, to date.

Nominated for Best Paper International Conference on Computer Vision (ICCV) 2023 website paper github
Artwork for Scale-MAE: A Scale-Aware Masked Autoencoder for Multiscale Geospatial Representation Learning
Emerging Technology and Policy Co-Design Considerations for the Safe and Transparent Use of Small Unmanned Aerial Systems

Ritwik Gupta, Alexander Bayen, Sarah Rohrschneider, Adrienne Fulk, Andrew Reddie, Sanjit A. Seshia, Shankar Sastry, Janet Napolitano

With the meteoric rise of small unmanned aerial systems, we discuss policy shortcomings in integrating sUAS technology in a safe fashion into our society. We suggest technology and policy co-design approaches to addressing these gaps in our systems.

Center for Security in Politics, UC Berkeley paper
Artwork for Emerging Technology and Policy Co-Design Considerations for the Safe and Transparent Use of Small Unmanned Aerial Systems
Region-level Active Detector Learning

Michael Laielli, Giscard Biamby, Dian Chen, Ritwik Gupta, Adam Loeffler, Phat Dat Nguyen, Ross Luo, Trevor Darrell, Sayna Ebrahimi

A new strategy that subsumes previous Image-level and Object-level approaches into a generalized, Region-level approach.

arXiv preprint paper
Artwork for Region-level Active Detector Learning
Creating xBD: A Dataset for Assessing Building Damage from Satellite Imagery

Ritwik Gupta, Bryce Goodman, Nirav Patel, Ricky Hosfelt, Sandra Sajeev, Eric Heim, Jigar Doshi, Keane Lucas, Howie Choset, Matthew Gaston

Preliminary work discussing xBD, the foundational dataset for assessing building damage after natural disasters from very-high resolution satellite imagery with over 850,000 annotations across 45,000 square kilometers.

Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) Workshops 2019 website paper github
Artwork for Creating xBD: A Dataset for Assessing Building Damage from Satellite Imagery
Open Problems in Robotic Anomaly Detection

Ritwik Gupta, Zachary T. Kurtz, Sebastian Scherer, Jonathon M. Smereka

Motivated by the development of ROS 2, this work discusses open problems in the field of robotic anomaly detection and presents an inverse reinforcement learning-based approach to detecting anomalous motion.

arXiv preprint paper
Artwork for Open Problems in Robotic Anomaly Detection