Research

OPERA: Automatic Offline Policy Evaluation with Re-weighted Aggregates of Multiple Estimators

Cohere

Offline policy evaluation (OPE) allows us to evaluate and estimate a new sequential decision-making policy's performance by leveraging historical interaction data collected from other policies.

Visit Site

Research Cohere

ResearchImproving Policy Learning via Language Dynamics DistillationCohere ResearchRewardBench 2: Advancing Reward Model EvaluationCohere ResearchSparkles: Unlocking Chats Across Multiple Images for Multimodal Instruction-Following ModelsCohere ResearchReality Check: A New Evaluation Ecosystem Is Necessary to Understand AI's Real World EffectsCohere ResearchThe Reality of AI and BioriskCohere ResearchThe Culture Funnel: You can’t align what isn’t in the dataCohere NewsAutomatic Testing for GraphQL APIsThenewstack NewsKyverno Defends Containers Against Security Configuration ErrorsThenewstack BlogsTailscale’s visual policy editor is in beta: add, edit, and deleteTailscale LearnNo-ETL: Query Raw CSV & JSON Files Directly with SQLMotherduck BlogsA decade of governance: Cloud Custodian at 10 and its role in the agentic AI eraCncf BlogsFrom data residency to digital sovereignty: Architectural patterns for cloud native platformsCncf BlogsImproving Composer through real-time RL · CursorCursor BlogsContinually improving our agent harness · CursorCursor ResourcesGet availability for multiple domains | Vercel SDKVercel Products & ServicesElevenLabs — Modern Slavery Policy StatementElevenlabs Products & ServicesNew this month: Performance upgrades, better LLM support, new blocks and moreGitbook BlogsWCAG 3.0’s Proposed Scoring Model: A Shift In Accessibility EvaluationSmashingmagazine ResourcesPolicy structFastly ResearchWhen Personalization Meets Reality: A Multi-Faceted Analysis of Personalized Preference LearningCohere