Loss Landscape Across the Double Descent Curve
Wolfram Community Editorial Board Staff Picks
This essay analyzes the relationship between model capacity and the flatness of minimizers across the double descent curve. Using filter-normalized loss surface visualizations, we find that the minimizer is the sharpest at the interpolation threshold and becomes gradually flatter as capacity moves in either direction—toward the under-parameterized sweet spot or deeper into the over-parameterized regime.