Take every open-weight model that discloses its parameter count, put total parameters on a log x-axis, put Terminal-Bench 2.1 score on the y-axis, and you get a reasonably tidy cloud sloping up and to the right. Bigger is better. This is the shape we have been trained to expect. Laguna S 2.1 scores 70.2%. To its right, at 1.6 trillion parameters, DeepSeek-V4-Pro-Max sits at 64.0. At 975B, Inkling is at 63.8. At 550B, Nemotron 3 Ultra is at 56.4. On DeepSWE, a harder and less saturated benchmark, the gap stops being subtle at all: Laguna S 2.1 scores 40.4 against DeepSeek-V4-Pro-Max's 9.0. A 13x parameter deficit paired with a 4x score advantage is the kind of result that usually means somebody broke the eval. Poolside seems to have anticipated that reaction, because they published every trajectory from every trial in the final evaluation set. You can go read what the model actually did. That decision tells you most of what you need to know about how this release was designed.
learn more