---

# Front Matter (YAML)

author: "contact@sebastienrousseau.com (Sebastien Rousseau)"
banner_alt: "An illustration representing a validation loss of 3.41, and a variance that swallows the result."
banner_height: "1280"
banner_width: "1920"
banner: "https://cloudcdn.pro/stocks/images/a-validation-loss-of-3-41-and-a-variance-that-swallows-the-result.webp"
cdn: "https://cloudcdn.pro"
charset: "UTF-8"
cname: "sebastienrousseau.com"
copyright: "© Copyright 2025 - 2026 - Sebastien Rousseau. All rights reserved."
date: "September 7, 2026"
description: "Router-S, a sparsely routed model, reaches a validation loss of 3.41 against a dense baseline trained on the same token budget. Each configuration was run three times, and the spread across those..."
format-detection: "telephone=no"
hreflang: "en"
icon: "https://cloudcdn.pro/clients/sebastienrousseau/v1/logos/sebastienrousseau.svg"
id: "https://sebastienrousseau.com/2026-09-07-a-validation-loss-of-3-41-and-a-variance-that-swallows-the-result"
image_alt: "Black and White Portrait of Sebastien Rousseau"
image_height: "162"
image_width: "162"
image: "https://cloudcdn.pro/stocks/images/sebastienrousseau.webp"
keywords: "validation, loss, 3.41, variance, swallows, result, router-s, dense, token, baseline, fixed"
language: "en-GB"
last_reviewed: "2026-09-07"
layout: "report"
locale: "en_GB"
logo_alt: "Logo for Sebastien Rousseau"
logo_height: "44"
logo_width: "44"
logo: "https://cloudcdn.pro/clients/sebastienrousseau/v1/logos/sebastienrousseau.svg"
menu: ""
measurementID: "G-169G4ET5HQ"
name: "Sebastien Rousseau"
permalink: "https://sebastienrousseau.com/2026-09-07-a-validation-loss-of-3-41-and-a-variance-that-swallows-the-result"
rating: "general"
referrer: "no-referrer"
robots: "index, follow"
schema: "FAQPage, Article"
seo_title: "A Validation Loss of 3.41, and a Variance That Swallows the Result"
short_name: "sebastienrousseau"
subtitle: "Router-S reaches a validation loss of 3.41 against a dense baseline on an identical token budget — but the same evaluation reports that seed variance exceeds the gap being measured, which is the more important number to sit with."
tags: "validation, loss, 3.41, variance, swallows, result, router-s, dense, token, baseline, fixed"
theme-color: "0, 83, 191"
title: "A Validation Loss of 3.41, and a Variance That Swallows the Result"
url: "https://sebastienrousseau.com/2026-09-07-a-validation-loss-of-3-41-and-a-variance-that-swallows-the-result"
viewport: "width=device-width, initial-scale=1, shrink-to-fit=no"

# RSS - The RSS feed front matter (YAML).
atom_link: "https://sebastienrousseau.com/2026-09-07-a-validation-loss-of-3-41-and-a-variance-that-swallows-the-result/rss.xml"
category: "AI"
docs: https://validator.w3.org/feed/docs/rss2.html
generator: "Static Site Generator (SSG) (version 0.0.26)"
draft_engine: "claude"
draft_model: "sonnet"
draft_version: "0.0.35-0.20260906233428-4200d43a9474+dirty"
item_description: "Router-S, a sparsely routed model, reaches a validation loss of 3.41 against a dense baseline trained on the same token budget. Each configuration was run three times, and the spread across those..."
item_guid: "https://sebastienrousseau.com/2026-09-07-a-validation-loss-of-3-41-and-a-variance-that-swallows-the-result/rss.xml"
item_link: "https://sebastienrousseau.com/2026-09-07-a-validation-loss-of-3-41-and-a-variance-that-swallows-the-result/rss.xml"
item_pub_date: "Mon, 07 Sep 2026 00:42:51 +0000"
item_title: "A Validation Loss of 3.41, and a Variance That Swallows the Result"
last_build_date: "Mon, 07 Sep 2026 00:42:51 +0000"
managing_editor: "contact@sebastienrousseau.com (Sebastien Rousseau)"
pub_date: "Mon, 07 Sep 2026 00:42:51 +0000"
ttl: "60"
type: "article"
webmaster: "contact@sebastienrousseau.com"

# Apple - The Apple front matter (YAML).
apple_mobile_web_app_orientations: "portrait"
apple_touch_icon_sizes: "192x192"
apple-mobile-web-app-capable: "yes"
apple-mobile-web-app-status-bar-inset: "black"
apple-mobile-web-app-status-bar-style: "black-translucent"
apple-mobile-web-app-title: "A Validation Loss of 3.41, and a Variance That Swallows the Result"
apple-touch-fullscreen: "yes"

# MS Application - The MS Application front matter (YAML).

msapplication-navbutton-color: "0, 83, 191"

# Twitter Card - The Twitter Card front matter (YAML).

twitter_card: "summary_large_image"
twitter_creator: "@wwdseb"
twitter_description: "Router-S, a sparsely routed model, reaches a validation loss of 3.41 against a dense baseline trained on the same token budget. Each configuration was run three times, and the spread across those..."
twitter_image: "https://cloudcdn.pro/clients/sebastienrousseau/v1/logos/sebastienrousseau.svg"
twitter_image_alt: "Logo of Sebastien Rousseau"
twitter_site: "@wwdseb"
twitter_title: "A Validation Loss of 3.41, and a Variance That Swallows the Result"
twitter_url: "https://sebastienrousseau.com/2026-09-07-a-validation-loss-of-3-41-and-a-variance-that-swallows-the-result"

excerpt: "Router-S reaches a validation loss of 3.41 against a dense baseline on an identical token budget — but the same evaluation reports that seed variance exceeds the gap being measured, which is the more important..."

# Humans.txt - The Humans.txt front matter (YAML).
author_website: "https://sebastienrousseau.com"
author_twitter: "@wwdseb"
author_location: "London, UK"
thanks: "Thanks for reading!"
site_last_updated: "2026-09-07"
site_standards: "HTML5, CSS3, RSS, Atom, JSON, XML, YAML, Markdown, TOML"
site_components: "Kaishi, Kaishi Builder, Kaishi CLI, Kaishi Templates, Kaishi Themes"

---

# A Validation Loss of 3.41, and a Variance That Swallows the Result

**Router-S reaches a validation loss of 3.41 against a dense baseline on an identical token budget — but the same evaluation reports that seed variance exceeds the gap being measured, which is the more important number to sit with.**

<!-- lead-start -->
<aside class="post-lead" aria-label="Article summary">
<p class="post-lead-tldr"><strong>TL;DR.</strong> Router-S, a sparsely routed model, reaches a validation loss of 3.41 against a dense baseline trained on the same token budget. Each configuration was run three times, and the spread across those seeds is larger than the gap the comparison is trying to detect — so the team reports the median rather than the mean, and the headline figure should be read as provisional rather than conclusive.</p>
<p class="post-lead-heading"><strong>Key takeaways</strong></p>
<ul class="post-lead-takeaways">
  <li><strong>Mechanism.</strong> Sparse routing reduces the compute a dense model spends on tokens that are trivially predictable.</li>
  <li><strong>Controls.</strong> The evaluation use fixes data order across runs and matches token budgets between Router-S and the dense baseline.</li>
  <li><strong>Result.</strong> Router-S reaches a validation loss of 3.41.</li>
  <li><strong>Caveat.</strong> Variance across three seeds per configuration exceeds the measured gap, so the median is reported instead of the mean.</li>
</ul>
</aside>
<!-- lead-end -->

> **Executive Summary**
>
> - Sparse routing is designed to cut the compute spent on tokens that are trivially predictable, rather than treating every token as equally costly to process.
> - Router-S and a dense baseline are compared on an identical token budget, which removes budget size as a confound.
> - Data order is held fixed across runs, so differences in loss cannot be attributed to shuffling.
> - Each configuration is trained three times, and the resulting seed variance exceeds the gap being measured — the reason the median, not the mean, is the reported statistic.
> - The headline figure, a validation loss of 3.41 for Router-S, needs to be read alongside that variance, not instead of it.

## What sparse routing is actually buying you

The premise behind Router-S is straightforward: a dense model spends the same amount of compute on every token, whether that token is genuinely hard to predict or almost free. Sparse routing changes that allocation. It reduces the compute a dense model spends on tokens that are trivially predictable, which means the parameters and computation saved there can, in principle, be redirected or simply not spent at all. This is a mechanism claim, not a hedge — it describes what the routing does, not what it might do. The interesting question is never whether routing changes the compute profile of a model; it obviously does by construction. The interesting question is whether that reallocation shows up as a measurable improvement once you control for everything else.

## Controlling for everything else is harder than it sounds

Two decisions in the evaluation design matter more than they might first appear. First, the use holds the data order fixed across runs, so that when loss moves, it cannot be explained away by a different shuffle exposing the model to easier or harder sequences in a different order. Second, Router-S is evaluated against the dense baseline on an identical token budget. That second control rules out the most common way sparse-versus-dense comparisons get muddied: giving one side more tokens to train on and calling the resulting gap an architectural win. With budget and ordering fixed, whatever difference remains between Router-S and the dense baseline has a narrower set of possible explanations.

That is a real methodological discipline, and it is worth taking seriously precisely because it makes the next problem harder to hide from.

## Three seeds, one honest admission

Each configuration is trained three times. That is not a large number of repeats, but it is enough to expose something the team did not have to disclose: seed variance exceeds the gap being measured. In other words, if you trained Router-S three times and the dense baseline three times, the spread of results within each of those triplets is wider than the difference between the two groups' central tendencies. That is a limitation claim, stated plainly, and it changes how the headline number should be read.

Because of that variance, the team reports the median across the three runs rather than the mean. This is a sensible response to a small, noisy sample — a median is less sensitive to a single outlying run dragging the reported figure in one direction — but it is also, itself, an admission. You do not reach for the median unless the mean would be telling a story the underlying runs do not fully support. Reporting practice here is doing some of the work that a larger seed count would otherwise do.

## Reading 3.41 for what it is

The headline result is that Router-S reaches a validation loss of 3.41. On its own, that is a specific, checkable number, produced under a fixed token budget and a fixed data order, against a dense baseline evaluated under the same conditions. That is worth stating plainly, because it is a demonstrated result, not a projection.

But it sits next to a second demonstrated result: the variance between seeds is larger than the gap the comparison is designed to detect. Put those two facts side by side and the honest reading is that 3.41 is a real, reproducible-in-principle measurement, but the comparison it is meant to support — is Router-S better than the dense baseline — has not yet cleared its own noise floor. Three runs per configuration is enough to notice that the noise floor exists. It is not obviously enough to say which side of it the true effect sits on.

None of this diminishes the value of the mechanism itself, or the discipline of the evaluation setup. Fixed token budgets and fixed data order are exactly the right controls to isolate an architectural effect from a training-recipe effect. What is missing is scale on the one axis that would resolve the remaining question: more seeds per configuration, run under the same fixed conditions, until the gap being measured is larger than the noise around it — or until it becomes clear that it isn't.

