<?xml version="1.0" encoding="utf-8" standalone="yes"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
  <channel>
    <title></title>
    <link>http://localhost:1313/blog/</link>
    <description>Recent content on </description>
    <generator>Hugo -- gohugo.io</generator>
    <language>en</language>
    <copyright>© 2026 </copyright>
    <lastBuildDate>Mon, 16 Feb 2026 00:00:00 +0000</lastBuildDate><atom:link href="http://localhost:1313/blog/index.xml" rel="self" type="application/rss+xml" />
    
    <item>
      <title>Data valuation for LLMs</title>
      <link>http://localhost:1313/blog/posts/metagrads/</link>
      <pubDate>Mon, 16 Feb 2026 00:00:00 +0000</pubDate>
      
      <guid>http://localhost:1313/blog/posts/metagrads/</guid>
      <description>&lt;h2 class=&#34;relative group&#34;&gt;Introduction&#xA;    &lt;div id=&#34;introduction&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#introduction&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;p&gt;One of the most critical components of training a foundation model (e.g., an LLM) is the choice of the training data.&#xA;A key challenge in designing the training data mix is estimating the value of a given data source.&lt;sup class=&#34;hovernote&#34; tabindex=&#34;0&#34; data-note=&#34;Example hover footnote: the value depends on how much the data improves downstream evaluation metrics.&#34;&gt;1&lt;/sup&gt;&#xA;&#xA;This challenge underlies questions such as:&lt;/p&gt;</description>
      
    </item>
    
  </channel>
</rss>
