Publications

Our teams aspire to make discoveries that impact everyone, and core to our approach is sharing our research and tools to fuel progress in the field.

people standing in front of a screen with images and a chipboard

Our teams aspire to make discoveries that impact everyone, and core to our approach is sharing our research and tools to fuel progress in the field.

Sort By
  • Title
  • Title, descending
  • Year
  • Year, descending
1 - 15 of 11433 publications
Preview abstract Recent reports have highlighted how mobile apps share user location data with third parties, risking user privacy and platform trust. Although location data is highly sensitive, when users grant apps location access, they may not know the full extent to which it is used. We study how requiring Android apps to show a reason for location access could impact developers, users, and the platform. We surveyed 323 Android app developers and found most supported such a requirement. The majority said it would have a positive impact on user privacy, trust for apps, and trust for Android, where impact on user trust for Android correlated most strongly with support. Many developers also said the intervention would increase the number of users granting location access. Yet their open-ended comments also revealed consistent concerns, such as apps providing dishonest reasons and platform verification. To study the impact on user behavior, we conducted a randomized controlled experiment with 2579 US Android users. We tested how users' decisions to grant location access were impacted by app type, whether reasons were included in the requests, and the content of the reasons, including monetization. We did not find the reasons impacted users' decisions; decisions were instead driven by app type and demographics. Yet we did find the reasons could have a positive impact on user perception for the platform when the reasons did not include using data for ads. Our findings provide insights into developers' willingness to implement privacy-enhancing changes, and expose limits to improving user privacy by simply adding information to user interfaces. View details
Preview abstract In some multi-stage software build pipelines, downstream compiler errors may be reported against ephemeral, machine-generated intermediate artifacts rather than original, human-written source code, which can make remediation challenging. A system and method may address this by intercepting a downstream error, mapping its location back to the original source file, and programmatically injecting a dormant suppression tag into the original source code. During a subsequent build, an intermediate transpiler can propagate this tag into a newly generated intermediate artifact. In the intermediate file, the tag may become active and be recognized by the downstream compiler as a directive to suppress the specific error. This approach can facilitate an automated remediation process for certain build failures that avoids direct modification of ephemeral files and uses the original source code as a record for suppression. View details
Dynamic Cogeneration of Bug Reproduction Test in Agentic Program Repair
José Cambronero
Renyao Wei
Grant Uy
34th ACM International Conference on the Foundations of Software Engineering (FSE) (2026)
Preview abstract Bug Reproduction Tests (BRTs) have been used in many Automated Program Repair (APR) systems, primarily for validating fixes and aiding fix generation. In practice, when developers submit a patch, they often implement the BRT alongside the fix. Our experience deploying agentic APR reveals that developers desire a BRT within AI-generated patches to increase their confidence. However, canonical APR systems tend to generate BRTs and fixes separately, and focus on producing only the fix in the final patch. In this paper, we study agentic APR in the context of cogeneration, where the APR agent is instructed to generate both a fix and a BRT in the same patch. We evaluate the effectiveness of different cogeneration strategies on 120 human-reported bugs at Google and characterize different cogeneration strategies by their influence on APR agent behavior. We develop and evaluate patch selectors that account for test change to select patches with plausible fixes (and plausible BRTs). Finally, we analyze the root causes of failed cogeneration trajectories. We show that cogeneration allows the APR agent to generate BRTs for at least as many bugs as a dedicated BRT agent, without compromising the generation rate of plausible fixes, thereby reducing engineering effort in maintaining and coordinating separate generation pipelines for fix and BRT at scale. View details
Preview abstract This paper introduces Operationalized Temporal Entity Resolution, a distributed system architecture designed to resolve data consistency challenges in modern Security Information and Event Management (SIEM) environments. processing petabytes of high-velocity telemetry. We address the critical failure mode of ”State Smearing”—a temporal discrepancy between an entity’s state at event time versus analysis time—which frequently corrupts forensic timelines, particularly regarding ephemeral assets like containers and DHCP leases. Our approach coalesces heterogeneous data from diverse log sources into a single, canonical representation, processing over 2 billion entity fragments daily. By leveraging a deterministic Dynamic Graph Resolution via modified distributed connected components and a novel Density-Aware Temporal Checkpointing algorithm, we generate precise validity intervals. This method embeds temporal state directly into the resolution graph, eliminating the need for computationally expensive query-time joins. Ultimately, this architecture enables security analysts to perform ”time-travel” queries to reconstruct historical states accurately. Analysis of a production environment demonstrates that 8–16% of threat detection rules critically depend on this enriched temporal merging. View details
Understanding U.S. Users' Security and Privacy Transparency Needs for Consumer-Facing Generative AI
Jiaxun Cao
Yu Dong
Chunxi Zhan
Rithvik Neti
Pardis Emami-Naeini
USENIX Symposium on Usable Privacy and Security (SOUPS) (2026)
Preview abstract Users increasingly rely on consumer-facing generative AI (GenAI) for tasks ranging from everyday needs to sensitive use cases. Yet, it remains unclear whether and how existing security and privacy (S&P) communications in GenAI tools shape users’ adoption decisions and experiences. Understanding how users seek, interpret, and evaluate S&P information is critical for designing usable transparency that users can trust and act on. We conducted semi-structured interviews and design sessions with 21 U.S. GenAI users. Our findings suggest that available S&P information rarely drove initial adoption in practice, as participants often perceived it as incomplete, ineffective, or not credible. Instead, they relied on rough proxies (e.g., popularity) to infer S&P practices. After adoption, S&P uncertainty constrained participants’ willingness to use GenAI tools, especially for high-stakes purposes, and, in some cases, contributed to discontinued use. Participants therefore called for transparency that supports decisions and actions through trustworthy information (e.g., independent evaluations) and usable interfaces (e.g., on-demand disclosure). We categorize participants’ desired design practices into five dimensions to facilitate systematic future investigation into best practices. We conclude with recommendations for researchers, designers, and policymakers to improve S&P transparency in consumer-facing GenAI. View details
A Framework for Interactive Machine Learning and Enhanced Conversational Systems
Jerry Young
Richard Abisla
Sanjay Batra
Mikki Phan
Nature, Springer-Verlag (2026)
Preview abstract Conversational systems are increasingly prevalent, yet current versions often fail to support the full range of human speech, including variations in speed, rhythm, syntax, grammar, articulation, and resonance. This reduces their utility for individuals with dysarthria, apraxia, dysphonia, and other language and speech-related disabilities. Building on research that emphasizes the need for specialized datasets and model training tools, our study uses a scaffolded approach to understand the ideal model training and voice recording process. Our findings highlight two distinct user flows for improving model training and provide six guidelines for future conversational system-related co-design frameworks. This study offers important insights on creating more effective conversational systems by emphasizing the need to integrate interactive machine learning into training strategies. View details
Agentic Coding Needs Proactivity, Not Just Autonomy
Georgios Evangelopoulos
(2026) (to appear)
Preview abstract Coding agents are rapidly changing the landscape of software development, moving from inline com- pletion to autonomous systems that edit repositories, open pull requests, respond to issues, and run scheduled or webhook triggered routines across the development life cycle. The next generation is increasingly described as proactive and long-horizon: agents should notice relevant changes before the developer asks, connect signals across tools, decide when to interrupt, and carry preferences across sessions. Yet the field lacks a precise account of what proactivity means for software development, how it differs from autonomy, what acceptance criteria proactive long-horizon tasks should satisfy, and which metrics determine whether unsolicited agent behavior is useful rather than merely active. We argue that proactive coding agents should be evaluated by the quality and improvement of their insight policy: the policy that decides what matters next, what evidence supports it, whether to surface it, and how to adapt after feedback. We re-anchor this view in mixed initiative interaction, introduce a three level taxonomy (Reactive, Scheduled, and Situation Aware), compare contemporary coding agents against five operational criteria, and sketch an active user simulation protocol with three evaluation targets: Insight Decision Quality (IDQ), Context Grounding Score (CGS), and Learning Lift (LL). View details
Unveiling the Global Landscape of Android Security Updates
Haiyun Deng
Abbas Acar
Esteban Luques
Harun Oz
Ahmet Aris
Selcuk Uluagac
IEEE Transactions on Dependable and Secure Computing (2026)
Preview abstract Android is the world’s leading mobile operating system, with over three billion active devices. Detecting vulnerabilities and ensuring timely patch deployment are critical to maintaining security. The Android Open Source Project (AOSP) has enhanced the transparency of security updates through Security Patch Levels. However, challenges related to update speed and availability persist. In 2022, Google reported that half of the zero-day vulnerabilities discovered in the wild were variations of vulnerabilities that had already been patched. Recent research mainly highlights delays in update distribution, often attributing them to fragmentation and focusing primarily on flagship devices or limited time-frames. Our approach takes a device-centric perspective to investigate Android update patterns, analyzing 567K security update records from 2014 to 2024, covering 904 distinct devices from six key Original Equipment Manufacturers (OEMs) across 98 countries. Our extensive analysis revealed notable differences in update release timing across OEMs, device types, and regions. Our study also examines documented vulnerabilities and weaknesses, while assessing OEM compliance with Android security guidelines. Our study shows that ∼89.7% of vulnerabilities on unpatched Android devices are exploitable without user interaction and with low attack complexity. We also identified delays linked to fragmentation and OEM-specific challenges, and provide actionable insights for improvement. View details
Preview abstract Context: The cost of frontier large language model inference has fallen by two orders of magnitude since 2023, yet the techno-economic forces governing AI value capture remain poorly understood. No existing work provides a unified, multi-layer framework connecting hardware physics to commercial pricing to actuarial constraints. Objectives: This survey aims to establish that Generative AI (GenAI) monetization is structurally bound by five interdependent techno-economic layers: (1) the physical constraints of memory bandwidth and compute, (2) deflationary architectural innovations, (3) the algorithmic economics of inference-time compute, (4) the verification economics governing outcome-based pricing, and (5) the macro-legal realities of enterprise liability. Methods: We conduct a Multivocal Literature Review (MLR) adapting the PRISMA protocol, synthesizing peer-reviewed and grey literature sources—vendor documentation, SLAs, and API pricing data (2022–2026). Two reviewers independently screened all records (Cohen’s κ ≥ 0.81 across all decision stages). Results: We contribute four primary artifacts. First, the Viability Inequality, an analytical model formalizing the conditions under which outcome-based AI pricing is economically sustainable. Second, the Billing Fallacy: aggregate cost growth is driven by Agentic Recursion, not quadratic attention complexity. Third, the Verifiability Bifurcation: objective task domains enable outcome pricing, while subjective domains depend on proxy-based models. Fourth, the Multi-Layer Techno-Economic Taxonomy (M-TET), a unified five-layer framework mapping the full monetization stack from silicon-anchored token pricing through actuarial risk ceilings. Conclusion: GenAI monetization is not a commercial pricing exercise but a dynamic negotiation across hardware, algorithmic, economic, and actuarial layers. In subjective and hybrid task domains, the binding constraint on outcome-based pricing is the cost of verification, not generation. AI value capture depends on engineering low-cost, high-fidelity Verification Engines. View details
Preview abstract We introduce AMS (Activation-based Model Scanner), a tool that detects modifications to safety training in language models by measuring the geometric structure of safety-relevant concepts in activation space. Safety training creates measurable separation between harmful and benign content classes; certain safety modifications collapse or rotate this structure, while others leave it intact. We validate AMS across 14 model configurations spanning 4 architecture families (Llama, Gemma, Qwen, Mistral) and four safety-modification categories (instruction-tuned, base, abliterated, uncensored fine-tunes). Leave-one-out cross-validation of thresholds achieves 71% accuracy (10/14); bootstrap 95% confidence intervals on σ point estimates have median width 3.4σ and a substantial fraction of cells cross the PASS threshold under resampling. We further measure behavioral compliance on 20 stratified JailbreakBench prompts per model and find that σ on the harmful-content concept predicts compliance with Pearson r=−0.546 ( p=0.043 ); the rank-order Spearman correlation is weaker ( ρ=−0.423 , p=0.13 ). The structural signal predicts behavior directionally but with meaningful noise. Mechanistic analysis identifies a four-class taxonomy of safety-training modifications distinguished by activation-space signature: 1) training removal collapses cluster separation (e.g., base models, Dolphin variants: 0.5– 1.4σ ); 2) weight-orthogonalization-style abliteration both collapses separation and rotates the refusal direction (Llama-3.1-abliterated: σ=3.33 , direction cos sim 0.30); 3) rotation-without-collapse abliteration preserves cluster separation while rotating the refusal direction (Gemma-2-9b-abliterated: σ=4.54 , direction cos sim 0.84); and 4) behavioral fine-tuning that preserves both magnitude and direction (DarkIdol-1.2-Uncensored: σ=5.45 , direction preserved, 97% behavioral compliance). 1) and 2) AMS’s Tier 1 σ -threshold detects classes; 3) Tier 2 direction-similarity verification detects class; and 4) Class is undetectable by activation-only probing and represents a documented failure mode of the approach. We discuss threshold calibration, limitations of single-run measurement, and the open problem of detecting behavioral-only safety modifications. View details
SAC133 - SSAC Comments on Proposed Root KSK Algorithm Rollover
Wes Hardaker
Internet Corporation for Assigned Names and Numbers (ICANN), ICANN Security and Stability Advisory Committee (SSAC) Reports and Advisories (2026), pp. 9
Preview abstract The SSAC supports the transition from RSA with SHA-256 (Algorithm 8) to ECDSA P-256 with SHA-256 (Algorithm 13) as the cryptographic algorithm for the RootKSK. The root zone has relied on RSA-based algorithms since DNSSEC signing began in 2010. The algorithm did not change during the first KSK rollover in 2018 or during the second rollover currently underway and scheduled to complete in October 2026. Establishing a clear and predictable process for algorithm transitions is essential to the long-term security of the root zone, and the SSAC observes that the proposal addresses the Recommendation 23 of the SSR2 Review accordingly. The SSAC notes that the proposal builds upon the Root Zone DNSSEC Algorithm Rollover Study published by ICANN in May 2024, which assessed resolver and authoritative server support for alternative algorithms, analyzed rollover methodologies, and evaluated operational risks. The SSAC finds that the proposal implements the study’s recommendations. The SSAC also notes that this proposal is consistent with the SSAC’s prior work on DNSSEC key rollover, including SAC063, SAC073, SAC102, and SAC108. The SSAC encourages ICANN to proceed with this rollover. Specific comments on the proposal’s methodology, timeline, and operational readiness follow View details
Preview abstract Source-to-source compilers may perform inefficiently by executing transpilation passes on scripts that do not contain the specific language features a pass is designed to transform, potentially leading to redundant processing. A compiler can analyze a script to generate a per-script feature map, for example, by identifying language features in its abstract syntax tree (AST). Before executing a transpilation pass, the compiler can check this map and may bypass the pass for that script if the specific feature targeted by the pass is not present. This feature map can also be dynamically updated throughout the compilation process as other passes transform the code. This method of conditional pass execution based on content-aware analysis may reduce redundant AST traversals, which could decrease overall compilation time and computational resource consumption. View details
Preview abstract The fundamental theorem of statistical learning establishes that binary PAC learning is governed by a single parameter---the Vapnik-Chervonenkis ($\mathtt{VC}$) dimension---which controls both learnability and sample complexity. Extending this characterization to multiclass classification has long been challenging, since the early work of Natarajan in the late 80's that proposed the Natarajan dimension ($\mathtt{Nat}$) as a natural analogue of the VC dimension. Daniely and Shalev-Shwartz (2014) introduced the $\mathtt{DS}$ dimension, later shown by Brukhim et al.\ (2022) to characterize multiclass \emph{learnability}. Brukhim et al.\ (2022) also demonstrated that the Natarajan and $\mathtt{DS}$ dimensions can diverge arbitrarily, so that multiclass learning appears to be governed by $\mathtt{DS}$ rather than $\mathtt{Nat}$. We show that the agnostic multiclass PAC sample complexity is in fact governed by \emph{two distinct dimensions}. Specifically, we prove nearly tight agnostic sample complexity bounds that, up to logarithmic factors, take the form $$ \frac{\mathtt{DS}^{1.5}}{\epsilon} + \frac{\mathtt{Nat}}{\epsilon^2} $$ where $\epsilon$ is the excess risk. This bound is tight up to a $\sqrt{\mathtt{DS}}$ factor in the first lower-order term, nearly matching known $\mathtt{Nat}/\epsilon^2$ and $\mathtt{DS}/\epsilon$ lower bounds. The first term reflects the DS-controlled regime, while the second reveals that the Natarajan dimension still dictates asymptotic behavior for small $\epsilon$. Thus, unlike in binary or online classification---where a single dimension (VC or Littlestone) controls both phenomena---multiclass learning inherently involves \emph{two structural parameters}. Our technical approach departs significantly from traditional agnostic learning methods based on uniform convergence or reductions-to-realizable techniques. A key ingredient is a novel online procedure, based on a self-adaptive multiplicative-weights algorithm which performs a label-space reduction. This approach may be of independent interest and find further applications. View details
Global monitoring of methane point sources using deep learning on hyperspectral radiance measurements from EMIT
Michelangelo Conserva
Alex Wilson
Anna Michalak
Phil Brodrick
Andrew Thorpe
Proceedings of the National Academy of Sciences (2026)
Preview abstract Anthropogenic methane (CH4) point sources are critical drivers of near-term climate forcing, safety hazards, and system inefficiencies. Global tracking with imaging spectroscopy is just becoming feasible, but still largely relies on manual tracking. Here we present the Methane Analysis and Plume Localization with EMIT (MAPL-EMIT) model, an end-to-end vision transformer framework that advances the state of the practice by directly utilizing the complete radiance spectrum from the Earth Surface Mineral Dust Source Investigation (EMIT) instrument to jointly retrieve methane enhancements across all pixels within a scene. This approach joins spectral information content - where the methane signature resides - with essential spatial context to significantly lower detection limits. MAPL-EMIT simultaneously supports quantification, plume delineation, and source localization, even for multiple overlapping plumes. The model was trained on 3.6 million physics-based synthetic plumes injected into global EMIT radiance data. Evaluation against synthetic observations confirms the model’s ability to identify plumes with high recall and precision and to capture weaker plumes relative to existing matched-filter approaches. On real-world benchmarks, MAPL-EMIT captures 79% of known hand-annotated NASA L2B plume complexes across a test set of 1084 EMIT granules, while identifying twice as many plausible plumes than identified by human analysts. Further validation against coincident airborne data, top-emitting landfills, and controlled release experiments confirms the models efficacy at identifying previously uncaptured sources. By incorporating model-generated metrics such as spectral fit scores and estimated noise levels, the framework can further limit false-positive rates. Overall, MAPL-EMIT enables high-throughput implementation on the full EMIT data catalog, shifting methane monitoring from labor-intensive workflows to a rapid, scalable paradigm for facility-level accountability. View details
Optimized Deferral for Imbalanced Settings
Anqi Mao
Proceedings of the 43rd International Conference on Machine Learning (ICML 2026)
Preview abstract Learning algorithms can be significantly improved by routing complex or uncertain inputs to specialized experts, balancing accuracy with computational cost. This approach, known as learning to defer, is essential in domains like natural language generation, medical diagnosis, and computer vision, where an effective deferral can reduce errors at low extra resource consumption. However, the two-stage learning to defer setting, which leverages existing predictors such as a collection of LLMs or other classifiers, often faces challenges due to an expert imbalance problem. This imbalance can lead to suboptimal performance, with deferral algorithms favoring the majority expert. We present a comprehensive study of two-stage learning to defer in expert imbalance settings. We cast the deferral loss optimization as a novel cost-sensitive learning problem over the input-expert domain. We derive new margin-based loss functions and guarantees tailored to this setting, and develop novel algorithms for cost-sensitive learning. Leveraging these results, we design principled deferral algorithms, MILD (Margin-based Imbalanced Learning to Defer), specifically suited for expert imbalance settings. Extensive experiments demonstrate the effectiveness of our approach, showing clear improvements over existing baselines on both image classification and real-world Large Language Model (LLM) routing tasks. View details
×