Technology ❯ Artificial Intelligence ❯ Model Evaluation ❯ Benchmarking

NoLima and NovelQA

New Papers Flag Retrieval as RAG’s Weak Link, Propose Practical Fixes

Fresh arXiv results highlight retrieval-induced failures in real datasets, offering open benchmarks plus methods that report measurable robustness gains.