lean4agent brings formal verification to agent workflows
a new framework uses the lean4 proof assistant to model and verify multi-step agent behavior, catching errors before execution.
topic
a new framework uses the lean4 proof assistant to model and verify multi-step agent behavior, catching errors before execution.
elmes* builds fine-grained rubrics to assess how large language models teach, not just what they know, across 330 long-tail educational scenarios.
new research shifts focus from behavior to internal mechanisms when assessing consciousness in animals and ai, finding current ai likely not conscious but leaving door open for insects and future machines.
gitco improves zero-shot forecast accuracy by filtering harmful patches from input context without retraining.
a new python package runs micropython inside a wasm sandbox for safe code execution with memory and cpu limits.
a field report on building a tiny woodland economy with qwen2.5-3b agents, showing how small models can drive emergent market behavior when paired with designed scarcity and sharp prompting.
a new python package helps scientists find governing equations from data by using structural skeletons and checking if parameters can be uniquely determined.
google's new agentic rag framework uses a sufficient context agent to iteratively search across data sources, improving accuracy on complex multi-hop questions.
a look at temperature scaling, platt scaling, and isotonic regression for fixing overconfident llms.
new framework decouples taming denominator from stochastic gradient noise to eliminate stationary bias in langevin algorithms.
a new clustering validation index called central description length uses probabilistic bounds on description length to evaluate clusters without labels, handling non-convex and irregular shapes better than traditional methods.
a study tests whether staged fractional-factorial experiments can identify stable early effects in micro-pretraining under tight compute budgets.
a stereological theory shows that standard llm benchmarks have a large structural blind spot, making rankings unreliable.
hello robot ships stretch 4, a safe home robot for research and disability assistance, focusing on real-world data collection over lab demos.
eva-bench data 2.0 adds itsm and healthcare hrsd domains, totaling 213 scenarios across 121 tools to test voice agents on realistic, domain-specific tasks.
new research finds emotional reliance on ai often starts incidentally during everyday tasks, not through dedicated companion apps, and can shift preferences away from human support.
a framework uses ontologies to generate test scenarios and issue verifiable trust certificates for enterprise ai agents before deployment.
new analysis of alternating power iteration for spiked tensor models gives finite-iteration error bounds and explains warm-start behavior without relying on specific initializations.