Moonshot AI is expanding its open technical work with PerceptionBench, a new evaluation set focused on basic visual perception. Its publication comes as FlashKDA receives renewed attention: the lab open-sourced this Kimi Delta Attention kernel implementation in April 2026, so it is not presented here as a July 28 release.
FlashKDA
The repository provides a CUTLASS-based implementation that can act as a backend for flash-linear-attention. Moonshot reports a 1.72x to 2.22x prefill speedup over a Triton baseline on H20 GPUs.
This is a targeted optimization, not a guaranteed acceleration for an entire application. The final result depends on hardware, context length, batch size and the rest of the inference stack.
PerceptionBench
The benchmark contains 3,000 verified questions derived from failures observed in frontier models. It organizes evaluation into ten capabilities including localization, counting, relations, depth, OCR and comparison.
The dataset and evaluation code are available so other laboratories can reproduce results. Like any benchmark, it may favour certain formats and does not replace testing on a product’s real images.
VERIFICATION SOURCEOfficial FlashKDA repository ↗VERIFICATION SOURCEPerceptionBench introduction ↗Does FlashKDA accelerate every model?
It is designed for architectures using Kimi Delta Attention and requires a compatible integration.
Does PerceptionBench measure general intelligence?
No. It targets atomic visual-perception capabilities while reducing the reasoning and external-knowledge burden.
Verified sources and original reporting.
Nexus AI wrote and contextualized this article using Moonshot AI · Apr and Jul 28, 2026. The complete story is on this page; the reference is provided so readers can check the original information.
Check the main source ↗
