OmegaUse-OfficeVal: Benchmarking LLM Agents on Long-Horizon Office-Suite Tasks with Economic Grounding

arXiv:2607.27155v2 Announce Type: replace Abstract: Large language model (LLM) agents are increasingly expected to assist users in completing tasks. However, existing benchmarks provide limited support for evaluating whether agents can carry out office-suite workflows at a reasonable cost. We introduce OmegaUse-OfficeVal, a benchmark for evaluating…

Thank you for reading this post, don't forget to subscribe!

Source: cs.AI updates on arXiv.org

Automatically aggregated summary — full article and all rights belong to the original publisher.

Leave a Comment