← All case studies

Correcting a severe hidden scaling defect in 4-bit quantized AI inference

We traced distorted outputs in a 4-bit quantized inference path to missing scale factors. A controlled component test confirmed the cause and demonstrated a host-side correction.

Project
Efficient local AI inference
Key outcomes
  • Isolated a scaling defect in a quantized expert-network computation.
  • Reduced relative L2 error to below 1% against the runtime’s quantized reference in the component test.
  • Demonstrated a host-side repair and specified proposed kernel changes for both prompt processing and token generation.

Background

Bringing a capable AI model onto local hardware means working within memory and response-time limits. Quantization can reduce memory use by representing values at lower precision. Evaluating the resulting model requires the runtime to apply that format and its scale factors correctly.

During our OCR work, we investigated a mixture-of-experts model, which routes token representations through selected expert networks. The affected expert-network path used four-bit floating-point values with higher-precision scale factors. Its runtime combined several operations in a kernel to execute that work efficiently.

Challenge

The optimized expert-network output differed sharply from the runtime’s quantized reference calculation. We needed to establish whether this was expected numerical variation or a defect in how the runtime executed the quantized operations.

The cause was hidden in how scale factors passed between operations. Activation quantization used a global factor for each expert, but a later fused calculation omitted that factor. Tests with all activation scale factors set to one could not expose the omission. An intermediate nonlinear operation then acted on values at the wrong scale. Multiplying the final output by a correction factor could not recover the intended calculation. The repair had to restore the scale at the correct point inside the execution path.

Solution

We followed the scale values through model conversion, the host interface, activation packing, and the fused calculation. A separate execution path that retained the factors helped establish the expected behavior.

We then compared the original path, an instrumented control, and a correction that folded the missing factors into coefficients supplied by the host. The comparison held the recorded inputs, expert routing, and source scale arrays constant. All measured paths used the same unmodified kernel archive, and independent processes repeated the experiment.

The corrected output was checked against the runtime’s quantized reference. We also built a reference that deliberately omitted the same factors. The original output closely matched this deliberately altered calculation. Together with a separate activation-packing check, these controls tied the error to the missing scales.

Result

In the controlled component test, the corrected path reached a cosine similarity of at least 0.99995 against the runtime’s quantized reference, with relative L2 error below 1%. The original path’s output had an L2 norm approximately 17,810 times that of the reference and a cosine similarity of approximately 0.371.

The demonstrated host-side correction restored close numerical agreement in the affected expert calculation. It provided a precise repair target for an execution fault that could distort later conclusions about OCR model quality.

We also specified proposed kernel patches for prompt processing and token generation, with explicit rules to prevent applying the scale twice. The measured component repair gave that work a defined target: restore the intended quantized calculation before drawing conclusions about the model’s recognition quality.

What would you like to achieve?

Tell us about your goal or the challenge in your way.

Let’s build a solution