Posts Tagged "dspark"

DeepSeek V4 Flash: when model, silicon and serving align

Under every AI product sits the same problem: turning GPUs and energy into tokens worth paying for, delivered fast enough for real work. DeepSeek V4 Flash is the cleanest alignment we have seen of a model built for serving efficiency, current silicon, and the engineering between. We measured what that alignment produces.

DeepSeek V4 Pro DSpark: the model isn't ready, the architecture is

We served DeepSeek V4 Pro for two days. What we wanted was hands on the architecture: attention that scales almost linearly instead of quadratically, and a speculative decoder that adapts to load. Both delivered. The preview checkpoint did not.