Research

The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI

Cohere

The race to train language models on vast, diverse, and inconsistently documented datasets has raised pressing concerns about the legal and ethical risks for practitioners.

Visit Site

Research Cohere

ResearchBridging the Data Provenance Gap Across Text, Speech, and VideoCohere ResearchMetadata Archaeology: Unearthing Data Subsets by Leveraging Training DynamicsCohere ResearchInvestigating Continual Pretraining in Large Language Models: Insights and ImplicationsCohere ResearchScalable Data Ablation Approximations for Language Models through Modular Training and MergingCohere ResearchBigScience: A Case Study in the Social Construction of a Multilingual Large Language ModelCohere ResearchSEACrowd: A Multilingual Multimodal Data Hub and Benchmark Suite for Southeast Asian LanguagesCohere EventsCloud Security Trends & Challenges: Complete GuideCybersecurity Exchange BlogsHow we grew Mintlify by doing things that don't scaleMintlify BlogsState of Speech: Our New Data Report Reveals ASR’s Untapped Potential - Deepgram Blog ⚡️Deepgram LearnBuild a Presentation Coaching Application with Recall - Deepgram Blog ⚡️Deepgram BlogsThe Most Important Work in AI Training Is Also the Most Overlooked - Deepgram Blog ⚡️Deepgram BlogsData Ingestion: 6 Ways to Speed Up Your ApplicationRedis BlogsHow to Conduct an AI Agent Security Audit [+ Checklist]Zapier BlogsHow to implement AI training for employeesZapier Blogs5 ways to automate Browse AIZapier BlogsHow to connect Google Sheets to NotionZapier NewsPancakes Are Delicious and Data Centers Are for Free StuffThenewstack NewsExplore and Visualize Data the Apache Superset WayThenewstack NewsTransform and Future-Proof Your Architecture with MACHThenewstack NewsSelecting the Right Database for Your MicroservicesThenewstack