prompt injection as role confusion
research shows llms prioritize text style over content, enabling jailbreaks through role confusion.
topic
research shows llms prioritize text style over content, enabling jailbreaks through role confusion.
a pipeline uses semantic retrieval and human review to measure how well a computer science program covers evolving curricular standards.
a new method extracts useful knowledge from large time-series models that don't fit scientific data, creating small, fast forecasters for sensor networks.
a new model shows llm agents have hidden internal beliefs that shape group decisions, explaining how confidence can exceed initial levels.
a new policy framework extends beyond permit/prohibit to handle obligations, waivers, and conflict resolution for llm-driven agents.
orbital ai data centers promise abundant solar power and cooling but must overcome radiation, heat rejection, and high maintenance costs.
aura iteratively learns human-consistency signals to audit llm-as-a-judge decisions with minimal human verification.
information lattice learning can be interpreted as a method for learning the structure of probabilistic graphical models by projecting probability distributions onto partition lattices and lifting rules back.
a new forecasting algorithm achieves optimal regret for both general and smooth proper losses, overcoming previous limitations in u-calibration.
mosaicleaks shows that deep research agents often expose private information when they search the web, and training them to be better at tasks makes the leakage worse.
navi-orbital runs a vision-language model on a low earth orbit satellite to classify scenes, describe content, and answer operator questions in plain english.
adding human collaborators to ai teams can hurt performance without structured coordination, but shared memory and approval gates help.
a new benchmark evaluates language model agents on long-horizon business management by simulating 500 days of running a startup.
google deepmind outlines a control roadmap to manage internal ai agents by treating them as potential insider threats and using layered monitoring.
a graph neural network study finds adding sparse station data to radar forecasts gives minimal gains, while numerical weather prediction and satellite inputs matter more.
a new mathematical framework connects shock-wave theory to the learning dynamics of stochastic gradient descent after removing parameter symmetries.
a new method trains task generators using a lightweight probe instead of costly solver rollouts, making it practical to create frontier tasks for reinforcement learning.
a statistical theory for offline policy optimization using only trajectory-level outcome labels instead of per-step rewards.