llms learn to lie consistently across models
a multi-model study shows that fine-tuned deceptive language models develop early, linearly detectable representations of dishonesty.
aisummaries filed under machine learning
a multi-model study shows that fine-tuned deceptive language models develop early, linearly detectable representations of dishonesty.
ainvidia releases cosmos 3, a single open model that generates video, reasons about physics, and predicts actions for robotics and autonomous systems.
ainew theory reveals when momentum helps or hurts in sparse training settings, based on two key timescales.
aia plain-language glossary of common ai terms like llm, rag, and hallucination for anyone who nodded along but wants clarity.
aia new method uses deep neural networks and adaptive prediction-powered learning to optimize treatment rules for bivariate survival outcomes in randomized trials.
aiuse python's textstat library to automatically detect overly complex language in entry-level job postings.
aia tutorial on using transformers.js for text classification, zero-shot labeling, and question answering directly in the browser with no server needed.
ailearn to read torch.profiler traces and tables to find bottlenecks in pytorch code, starting with a simple matrix multiply and add.
aia compact binary mask reverses most knowledge edits in language models, revealing a shared mechanism behind diverse factual updates.
aia new mirror-prox temporal-difference method uses behavior-policy transition information instead of feature covariance to speed up off-policy prediction.
ainew lower bounds show that the bandwidth term in federated probe-logit distillation is tight, and the method extends to nodes with different upload budgets.
aia new llm-based agent called trace uses tool planning to optimize drug-like molecules over multiple steps, improving properties while keeping key structures intact.
aireplacing the auxiliary covariance matrix with the behavior bellman matrix improves stability in off-policy temporal-difference learning.
aiitbench-aa evaluates ai agents on kubernetes incident response, with claude opus 4.7 leading at 47% accuracy.
aigoogle research highlights from i/o 2026 include new ai tools for scientific discovery, health coaching, and edge computing, plus advances in weather prediction and model factuality.
aibuild practical ai assistants for job search, research, invoice processing, and more with step-by-step guides.
ailearn to fine-tune local language model parameters, optimize hardware, and format prompts using ollama's modelfile, environment variables, and go templates.
aia new method extends schrödinger bridge models for time series by using a frozen, triangular reference process across latent volatility levels, preserving the h-transform structure even with degenerate covariance.
aia survey examines how mixture-of-experts methods address key multimodal learning issues like scalability, representation, and fusion.
aia foundation model downscales global ai weather forecasts from 28 km to 1 km resolution, producing hourly 67-hour forecasts of eight surface variables.
ai