eva-bench grows to three enterprise voice agent domains
eva-bench data 2.0 adds itsm and healthcare hrsd domains, totaling 213 scenarios across 121 tools to test voice agents on realistic, domain-specific tasks.
aisummaries filed under research
eva-bench data 2.0 adds itsm and healthcare hrsd domains, totaling 213 scenarios across 121 tools to test voice agents on realistic, domain-specific tasks.
ainew research finds emotional reliance on ai often starts incidentally during everyday tasks, not through dedicated companion apps, and can shift preferences away from human support.
aia framework uses ontologies to generate test scenarios and issue verifiable trust certificates for enterprise ai agents before deployment.
ainew analysis of alternating power iteration for spiked tensor models gives finite-iteration error bounds and explains warm-start behavior without relying on specific initializations.
aithe ieee p3109 draft standard defines parameterized binary floating-point formats and operations to support efficient, consistent machine learning computations.
aia semiotic scaffolding called peel reveals systematic distortions in ai-generated text condensations, showing fluency does not equal fidelity.
aia large-scale analysis across four eeg datasets examines how different scalp regions contribute to predicting cognitive workload, revealing consistent patterns and practical implications for sensor selection.
aia guide to five foundational papers covering transformer architecture, few-shot learning, scaling laws, instruction tuning, and retrieval-augmented generation for understanding large language models.
aia new benchmark uses public prediction-market and blockchain data to evaluate how well models predict individual beliefs and trades.
aigoogle research releases its ai flood forecasting framework on github, enabling national agencies to train and customize models with local data.
aidirect preference optimization reduced text degeneration by an average of 59.4% across five ocr model families by using the model's own failure outputs as rejection pairs.
aia new framework lets human agents approve or reject algorithmic price suggestions, using old pricing data to skip the slow start typical in sparse booking markets.
aia new method reduces the cubic complexity of gaussian processes with gradients by using exact gradient reduction and vecchia approximation.
aia position paper argues that in high-dimensional settings, many different mechanisms can produce the same data, so predictive success does not prove a model has found the true mechanism, and large language models can hide this by giving a single fluent explanation.
aiaura-mem uses a learned gate to write only when observations change actions, keeping memory fixed at 4,224 bytes regardless of episode length.
aiperiodic and soft target updates can guarantee convergence in linear q-learning under explicit spectral and step-size conditions.
aia new method uses gradient tests instead of validation loss to decide when to stop training gradient boosted trees, avoiding the need for a patience parameter.
aia new diagnostic reveals that common anomaly detection benchmarks become unreliable when held-out classes overlap with normal data in representation space.
aia new framework aligns structured electronic health record representations with large language models to improve clinical prediction and reasoning.
aimicrosoft announced two new language models, mai-thinking-1 and mai-code-1-flash, with low active parameter counts and claims of clean training data.
ai