“The Chinese AI lab just dropped V4.1-Flash, a leaner model that slashes agent running costs. By cutting back on compute and memory compared to the previous version, it lets teams tackle the same workloads without racking up massive API bills ($0.30 per million input tokens). On OpenDesign benchmarks, it reached 98% of”
The Chinese AI lab just dropped V4.1-Flash, a leaner model that slashes agent running costs. By cutting back on compute and memory compared to the previous version, it lets teams tackle the same workloads without racking up massive API bills ($0.30 per million input tokens). On OpenDesign benchmarks, it reached 98% of






