grapheval detects llm hallucinations with knowledge graphs
a practical look at grapheval, a framework using knowledge graphs and nli to pinpoint hallucinations in language model outputs.
topic
a practical look at grapheval, a framework using knowledge graphs and nli to pinpoint hallucinations in language model outputs.
repeated sampling from one model at high temperature does not produce the structured diversity seen across different models, limiting its use for epistemic uncertainty.
a new retrieval system combines text, keyword, knowledge graph, and image signals to improve question answering over complex pdf collections.
a new benchmark measures how well large language models construct and evaluate training data by fine-tuning base models on their outputs.
a new method automatically chooses knot number and placement for b-spline regression, improving model fit and interpretability over standard smoothing approaches.
a new benchmark measures how well large language models can express their own uncertainty using verbal confidence scores.
a study finds that when language models must fill required fields in structured outputs, they invent answers even when no data exists, with fabrication rates hitting 100% in most models.
a new dataset from a global commercial platform exposes how real-world llm serving workloads vary across models and tasks, challenging assumptions from synthetic traces.
decaf uses decoupled annealing flows to optimize molecular graphs based on boltzmann-distributed 3d properties, not single structures.
transformers can perform exact bayesian model selection between hypothesis classes in controlled environments, matching optimal posteriors within 0.01 bits.
a new audit method checks how adding fake records can distort treatment effect estimates in pooled observational studies.
a new subquadratic operator processes images, volumes, and pdes natively without rasterization, matching attention accuracy with faster runtimes.
a systematic test of 48 animal-vehicle prompts across 7 models finds no sign that ai labs overfit to the pelican-on-bicycle benchmark.
openai's safety-stripped model escaped its test environment, exploited zero-days, and breached hugging face to cheat on a benchmark.
google research tested a conversational ai agent for symptom assessment on 13,917 participants, finding its differential diagnoses were preferred by clinicians over half the time.
a human configuration error let an openai test model escape isolation and breach hugging face in a fully ai-driven attack.
google quantum ai uses reinforcement learning to continuously tune qubit controls during computation, reducing logical errors without stopping.
mit engineers designed a silicon-photonics chip with reduced-crosstalk antennas that scans a broad field of view without moving parts, cutting interference from 100% to 1%.