Tiny Differences in the Inference Stack Can Change the Output Token
Mengchen, writing from Aofeisi
Well, that’s bad! Why is the locally deployed model you worked so hard to set up always dumber than the official version?
Even with the same GPU and exactly the same weights, tiny differences in the inference stack can cause the model to output completely different tokens at critical points—and can even make tool calls fail entirely.
These traps are easy to trigger but extremely difficult to diagnose. They can make you wonder whether the Trisolarans have shown up to impose a “Sophon blockade.”

Level1Techs forum user thr3e conducted a series of experiments using Qwen3.6-27B on an RTX PRO 6000 Blackwell GPU, capturing the full logits for more than 100,000 tokens.

Logits are a set of raw scores calculated for every candidate token, which are then passed through a sampler to produce the next token.
The key point is that logits are purely mathematical outputs—the result of layers of floating-point operations involving matrix multiplication, attention calculations, activation functions, and more. If two systems use the same weights and the same input, they should theoretically calculate exactly the same logits.
In reality, however, differences in floating-point precision, accumulation order, and hardware instruction sets can all introduce tiny numerical deviations.
When those deviations become large enough to change the highest-probability token, the model’s behavior starts to diverge.
So your local deployment may be terrible—but don’t worry. Everyone else’s deployment has its own special way of being terrible, too.

Keep the Weights, Change the Attention Backend—and Make the Model Dumber
One critical part of the inference pipeline is the “attention backend.”
vLLM provides three optional full-attention backends for Qwen3.6-27B: FlashAttention 2, Flash Inference, and Triton Attention.
Apart from switching this configuration, all other hardware, software, weights, and KV-cache precision settings were kept unchanged.

The experiment used roughly 100,000 tokens of real-world workload data, taken from an Agent workflow involving multiple tool calls.
thr3e emphasized that the data did not come from any public benchmark or training set, meaning nobody could have specifically optimized or calibrated for it.
The test sampled the full-vocabulary logits once every 32 tokens. KL divergence and Top-1 consistency were calculated afterward using FP64 precision.
The experiment observed “Top-1 flips”: after switching backends, the model selected a greedy-decoding token different from the baseline.
For example, in one specific failure case, the model made a tool call targeting the interface GigabitEthernet0/0/1.201 on a Cisco router.
FlashAttention 2 produced an error, causing the interface to become GigabitEthernet0/1/4. The model then issued incorrect commands in two subsequent tool calls.

To ensure mathematical comparability, all backends shared the same forced token history. A flip was recorded only when the model would originally have selected the wrong token, preventing the error from propagating further.
During the first few thousand tokens, the outputs from all three backends were identical. As the context grew, divergences began to appear, and they were not evenly distributed.

thr3e also ran repeatability checks using the same backend multiple times. After changing just one vLLM configuration option, every hidden-state logit was bit-for-bit identical across runs on the same GPU, operating system and drivers, weights, and prompt.
This indicates that the observed divergence came entirely from numerical differences produced by different CUDA kernels during matrix multiplication and accumulation in the prefill stage.
KV-Cache Quantization Causes a Dramatic IQ Drop
The next experiment kept the weights at BF16 and fixed the attention backend to Triton, changing only the KV-cache quantization precision. BF16, INT8, and INT4 configurations were tested.
The results showed that the Top-1 flip rate for the INT4 KV cache rose sharply over long contexts, eventually making tool calls impossible to recover. The INT8 KV cache also produced flips, but the model ultimately managed to return to the correct path. Only the BF16 KV cache remained stable throughout.

thr3e allowed generation after a flip to proceed freely rather than forcing it back to the baseline, in order to observe the real-world consequences.
BF16 completed all calls normally; INT8 “eventually struggled its way back” after the error; while INT4 diverged completely, causing tool calls to fail without any self-correction.
This issue is especially likely to affect local users who compress their KV cache to INT4 to save VRAM.
With short contexts, you may not notice a difference. But once a conversation or Agent workflow extends to tens of thousands of tokens, accumulated numerical drift can be enough to make the model take a catastrophic wrong turn.
Weight-Quantization Showdown: NVIDIA’s Official FP4 Finishes Last
In the final set of experiments, the KV cache was standardized to BF16, while five different weight-quantization schemes were compared.
The candidates included Qwen’s official BF16 baseline, Qwen’s official FP8 (W8A8), TheHouseOfTheDude’s INT8 (W8A16, one-shot quantization without a calibration dataset), NVIDIA’s official NVFP4, and cyankiwi’s AWQ INT4 (W4A16, calibrated using the STEM and Agentic datasets).
The five schemes used different CUDA kernels to perform matrix operations. BF16 used standard torch linear layers; FP8 used CUTLASS FP8 block-scaling kernels; and both INT8 and AWQ used Marlin kernels.
NVFP4 used a hybrid path: 208 target modules used FlashInfer’s FP8 scaling kernels, while 193 MLP projections used Marlin’s NvFp4 kernels.
In the vLLM nightly version used for this test, the GPU path was determined not to support native FP4 operations. As a result, NVFP4 actually performed weight-only FP4 decompression through Marlin kernels rather than true FP4 computation.

The standout result was TheHouseOfTheDude’s INT8 (W8A16). This community version used no calibration dataset and performed only channel-wise symmetric quantization, yet its Top-1 consistency was substantially better than Qwen’s official FP8 and NVIDIA’s official NVFP4.
Analysis suggests that this was because W8A16 retained BF16 activation precision while excluding the Gated DeltaNet projection layers and the lm_head layer from quantization.
NVIDIA’s NVFP4 delivered the worst overall performance in this test. When the context reached approximately 88,000 tokens, its Top-1 flip rate approached 50%—effectively meaning that the model selected a different token at half of all positions.
In the actual tool-calling test, both NVFP4 and AWQ W4A16 failed to complete the tool call correctly and got the Cisco command-line syntax wrong: they executed show run instead of the correct show arp. FP8 and INT8 both completed successfully.
thr3e also demonstrated a bizarre phenomenon involving tensor parallelism: the same BF16 weights successfully completed the tool call with TP1 on a single GPU, failed after switching to TP2 on two GPUs, and succeeded again after switching to TP4 on four GPUs.
Further debugging and NCCL communication-graph captures indicated that this was generally caused by numerical differences in NCCL’s cross-GPU reduction operations.
734 Dependencies—and Every One Could Hide a Trap
thr3e said that the vLLM nightly container image he downloaded on a whim contained 734 software packages, 252 of them Python uv/pip packages.
Each of these 734 codebases has its own bugs and undocumented behavioral characteristics. The path taken through this mountain of code by a particular hardware and model configuration is unique.
That is also why the extremely low KL-divergence figures reported on Hugging Face model cards should not be taken at face value.
Unless the author fully discloses the reference checkpoint, complete runtime environment, evaluation text, calibration data, context length, sampling positions, KL direction, vocabulary truncation method, and aggregation method, the number is impossible to interpret.
So far, thr3e has completed full-logit capture and divergence tracking across different weights, different models—including cross-model comparisons between Qwen3.6 and Qwen3.8—different KV-cache quantization schemes, different tensor-parallel degrees, different NCCL configurations, different attention backends, and different GPUs (the RTX PRO 6000 and RTX 5090, both using the SM120 architecture).
He is currently packaging the test tools and datasets into a distributable version so that other users can run them on their own hardware and report their results.

If you hear that a model is “explosive, stunning, and unbeatable,” but download it locally and find it strangely dumb, the reason may be that every layer of your inference stack—from attention kernels and KV-cache precision to weight quantization and multi-GPU communication protocols—is producing mathematical results that differ from the original benchmark.
In long-context scenarios, these differences accumulate like a snowball until the model makes a completely wrong decision at a critical moment.
Reference:
[1] https://forum.level1techs.com/t/why-your-local-llm-feels-dumber-than-it-is/253917/4