Resources / References / 01

Four to One: Auditing the Leaked DeepSeek Notes

Ryan Cunningham · Published August 6, 2026

E² review · September 25, 2026

Separate the claim, its evidence and the check performed.

Numerical checks have moved to the Audit tab.

Claims ledger

Each record separates the claim’s type, its support and the check actually performed. “Reported fact” describes an attributable statement; it does not certify the underlying event. Read the rubric.

Central claims, not an exhaustive audit of every sentence. Original footnote links are preserved in Body; only sources named in each review record were checked to the extent stated.

Symbol key

Symbols supplement the words; they do not replace them.

Reported source Assumption Calculation Measurement

Source checked Arithmetic reproduced Conflict Check still open

Support uses one to five filled marks, from conjectural to established within scope. This is an ordinal evidence scale, not a probability. A dash means the scale does not apply.

  1. F01

    DeepSeek received about 16,000 Ascend 950 chips.

    Type
    Alleged leak
    Support
    Reported
    Review
    Not independently verified

    The article attributes this allocation to reporting about leaked remarks. Neither the original meeting record nor delivery records were authenticated in this review. Repetition of the report is not independent corroboration.

    Evidence and revision conditions

    Passage in the article

    What would change this assessment: An authenticated primary account or delivery record identifying date, quantity and chip variant.

    E² editorial review · September 25, 2026 · Revision 1

  2. F02

    The allocation is 950DT hardware in two Atlas 950 SuperPoDs.

    Type
    Model assumption
    Support
    Not applicable
    Review
    Source checked

    The author marks the topology as a guess. Huawei’s roadmap schedules 950DT and Atlas 950 for Q4 2026, after the article’s August publication. An early allocation is possible; advertised maximum size does not establish the delivered SKU or topology.

    Evidence and revision conditions

    Passage in the article

    Depends on F01, F03.

    What would change this assessment: A deployment disclosure specifying SKU, system layout and operational date.

    E² editorial review · September 25, 2026 · Revision 1

  3. F03

    Huawei advertised an Atlas 950 system with up to 8,192 NPUs.

    Type
    Reported fact
    Support
    Supported
    Review
    Source checked

    The vendor announcement supports this narrow statement about advertised scale. It does not measure workload throughput or verify DeepSeek’s allocation.

    Evidence and revision conditions

    Passage in the article

    What would change this assessment: A revised vendor specification. Actual performance needs separate workload measurements.

    E² editorial review · September 25, 2026 · Revision 1

  4. F04

    The LongCat model card reports about 48B active parameters and more than 35T training tokens.

    Type
    Reported fact
    Support
    Supported
    Review
    Source checked

    The first-party card supports the reported architecture and corpus scale. The 35T scenario uses a rounded lower input, not the exact disclosed training total. The underlying training logs were not independently measured.

    Evidence and revision conditions

    Passage in the article

    What would change this assessment: Versioned architecture details or logs that revise the active count or exact token total.

    E² editorial review · September 25, 2026 · Revision 1

  5. F05

    At 48B active parameters and 35T tokens, the 6N training estimate is 1.008 × 10²⁵ FLOPs.

    Type
    Calculated output
    Support
    Not applicable
    Review
    Arithmetic reproduced

    6 × 48 × 10⁹ × 35 × 10¹² = 1.008 × 10²⁵. This is a conditional estimate of model FLOPs, not measured device work. Efficiency must use the same FLOP convention to avoid double-counting overhead.

    Evidence and revision conditions

    Passage in the article

    Depends on F04.

    What would change this assessment: Changed inputs, or an explicit architecture-specific accounting with a consistently defined efficiency denominator.

    E² editorial review · September 25, 2026 · Revision 1

  6. F06

    The printed LongCat base scenario takes 32.41 days on 50,000 chips.

    Type
    Calculated output
    Support
    Not applicable
    Review
    Arithmetic reproduced

    At 400 trillion FLOP/s per chip, 30% MFU and 60% goodput, the result is 1.620 million accelerator-days. This does not identify the historical chip SKU or actual duration. A numerical TFLOP/s input requires a factor of 10¹² in the denominator.

    Evidence and revision conditions

    Passage in the article

    Depends on F05, F08.

    What would change this assessment: A different dense peak, efficiency denominator, training volume or measured run duration.

    E² editorial review · September 25, 2026 · Revision 1

  7. F07

    The LongCat caption’s 22–98-day range follows from fixed 30% MFU and 40–75% goodput.

    Type
    Calculated output
    Support
    Not applicable
    Review
    Unresolved conflict

    Those fixed-MFU inputs yield 25.93–48.61 days. The wider 22.22–97.22-day range requires varying MFU from 15% to 35% as well. The caption and Appendix A describe different sweeps; neither range is a confidence interval.

    Evidence and revision conditions

    Passage in the article

    Depends on F05.

    What would change this assessment: Correct the caption to name both varying inputs, or replace its range with the fixed-MFU result.

    E² editorial review · September 25, 2026 · Revision 1

  8. F08

    Ascend goodput is 40–75% across the modeled training scenarios.

    Type
    Model assumption
    Support
    Not applicable
    Review
    Source checked

    The author explicitly calls this weakly sourced. It is a sensitivity range, not a measured fleet statistic. Holding it and MFU constant when scaling from thousands to hundreds of thousands of chips is a further assumption.

    Evidence and revision conditions

    Passage in the article

    What would change this assessment: Measured productive wall-clock fractions by fleet size, topology, workload and restart policy.

    E² editorial review · September 25, 2026 · Revision 1

  9. F09

    Pangu Ultra MoE’s authors report 30.0% MFU on 6,000 Ascend NPUs.

    Type
    Reported measurement
    Support
    Supported
    Review
    Source checked

    The abstract supports this first-party measurement report. It anchors one workload and scale; it does not independently validate the article’s broader Ascend fleet scenarios.

    Evidence and revision conditions

    Passage in the article

    What would change this assessment: Independent reproduction, or detailed counterevidence about the workload and MFU denominator.

    E² editorial review · September 25, 2026 · Revision 1

  10. F10

    Kimi K3’s model card reports 104B active parameters.

    Type
    Reported fact
    Support
    Supported
    Review
    Source checked

    The model summary reports this architecture figure. Combined with an assumed 35T tokens, the 6N calculation yields 2.184 × 10²⁵ FLOPs. The assumed corpus does not become a reported training fact.

    Evidence and revision conditions

    Passage in the article

    What would change this assessment: A revised model card or a disclosed training-token total.

    E² editorial review · September 25, 2026 · Revision 1

  11. F11

    Kimi K3 was most likely trained on foreign silicon.

    Type
    Inference
    Support
    Reported
    Review
    Not independently verified

    The article cites reporting about a Hopper cluster. A plausible 47-day modeled schedule alone cannot establish which machines trained the model. The cited historical reporting was not authenticated in this audit.

    Evidence and revision conditions

    Passage in the article

    Depends on F10.

    What would change this assessment: Primary training provenance, allocation records or a technical report naming the fleet.

    E² editorial review · September 25, 2026 · Revision 1

  12. F12

    50,000 GB300s take one quarter of the cited 99-day Ascend run.

    Type
    Calculated output
    Support
    Not applicable
    Review
    Unresolved conflict

    The body gives 53 days, so its time ratio is 53/99 = 0.535. One quarter would be 24.75 days. Confirm model version, precision, efficiency and GPU-versus-superchip units before choosing a replacement headline.

    Evidence and revision conditions

    Passage in the article

    What would change this assessment: A consistent configuration export that reconciles the summary and body.

    E² editorial review · September 25, 2026 · Revision 1

  13. F13

    The 33T Flash / 8.6T Pro headline denotes generated tokens per day.

    Type
    Calculated output
    Support
    Not applicable
    Review
    Contradicted as labeled

    At the printed 16:1 input:output mix, 1.9T Flash output implies 32.3T total; 0.5T Pro output implies 8.5T total. The larger numbers include input tokens. “Generated” is the wrong label under those inputs.

    Evidence and revision conditions

    Passage in the article

    What would change this assessment: Relabel total tokens and keep generated/output tokens separate throughout.

    E² editorial review · September 25, 2026 · Revision 1

  14. F14

    The Pro body’s roughly 14T total is the same 16:1 scenario as the opening.

    Type
    Calculated output
    Support
    Not applicable
    Review
    Unresolved conflict

    At 0.5T output, 14T total instead matches the 26.8:1 Pro mix in footnote 39: 0.5 × 27.8 = 13.9. Changing input mix also changes prefill work; multiplying a fixed output projection is only a reconciliation of labels, not a rerun.

    Evidence and revision conditions

    Passage in the article

    Depends on F13.

    What would change this assessment: Name the workload mix for each result and recompute sustained throughput at that mix.

    E² editorial review · September 25, 2026 · Revision 1

  15. F15

    Overclock reproduces Tensor Economics calculations within 2–5%.

    Type
    Calculated comparison
    Support
    Reported
    Review
    Not reproduced

    This is author-reported implementation agreement. Even if reproduced, agreement with another model is not independent measurement of deployment performance.

    Evidence and revision conditions

    Passage in the article

    What would change this assessment: Publish reference cases, inputs, expected values and executable comparison code.

    E² editorial review · September 25, 2026 · Revision 1

  16. F16

    Across 44 InferenceX runs, Overclock overpredicts throughput by median 1.81×, with p10–p90 of 1.39–2.18×.

    Type
    Comparison with independent measurements
    Support
    Reported
    Review
    Not reproduced

    The author says these measurements were held out. The public historical query is retained with this audit, but the exact 44-row selection and paired Overclock predictions are unavailable here. We cannot recompute the median, quantiles or flatness by concurrency from aggregate statements.

    Evidence and revision conditions

    Passage in the article

    What would change this assessment: Publish result IDs, paired predictions, metric units, versions, exclusions and the tuning/holdout split.

    E² editorial review · September 25, 2026 · Revision 1

  17. F17

    The stated fleet capacities meet 20 output tokens per second per user.

    Type
    Calculated output
    Support
    Not applicable
    Review
    Not validated

    The model selects its batch using predicted interactivity. The article also reports a 2.8× median overestimate of that metric. Dividing a projected 20 by 2.8 gives 7.14 as an illustrative sensitivity, not a pointwise correction. Aggregate throughput derating cannot validate the batch-selection constraint.

    Evidence and revision conditions

    Passage in the article

    Depends on F02, F16, F18.

    What would change this assessment: Recalibrate latency, rerun the batch sweep, and validate the selected configurations against measured service targets.

    E² editorial review · September 25, 2026 · Revision 1

  18. F18

    Overclock’s per-user interactivity is 2.8× high at the median.

    Type
    Comparison with independent measurements
    Support
    Reported
    Review
    Not reproduced

    Disclosing this residual is a strength. The figure still requires paired data and an explicit latency statistic. Its effect on the selected operating point matters more than whether the aggregate-throughput error looks stable.

    Evidence and revision conditions

    Passage in the article

    Depends on F16.

    What would change this assessment: Publish the per-row interactivity residuals and evaluate held-out configurations after recalibration.

    E² editorial review · September 25, 2026 · Revision 1

  19. F19

    Cross-chip and cross-model ratios remain sound despite absolute throughput errors.

    Type
    Inference
    Support
    Conjectural
    Review
    Not established

    Stable bias on B200/B300 does not establish matched bias on Ascend superpods or other models. If two biases independently occupy 1.4–2.2, a predicted ratio can be 0.64–1.57 times the true ratio. This is a sensitivity bound, not an empirical interval.

    Evidence and revision conditions

    Passage in the article

    Depends on F16.

    What would change this assessment: Matched measurements across the compared systems, workload mixes and precisions, with ratio residuals reported.

    E² editorial review · September 25, 2026 · Revision 1

  20. F20

    The 16,000-chip fleet comfortably covers total DeepSeek demand.

    Type
    Inference
    Support
    Conjectural
    Review
    Not established

    Flash capacity is compared with mixed-model demand. Router traffic is a subset; revenue-based estimates omit free usage. Price, cache, workload mix and peak load also matter. The printed figures do not establish enough capacity for total demand at the service target.

    Evidence and revision conditions

    Passage in the article

    Depends on F01, F02, F13, F14, F17.

    What would change this assessment: A matched output/input workload mix, full demand coverage, measured sustainable capacity and peak-load margin.

    E² editorial review · September 25, 2026 · Revision 1

  21. F21

    Nine months is a ceiling beyond which a training run is not viable.

    Type
    Inference
    Support
    Reported
    Review
    Unresolved interpretation

    Epoch estimates an economic optimum near 8.6 months with a 6.1–14.2-month interval under its assumptions. It is not a physical ceiling. Restricted access to future hardware can change the opportunity cost of waiting.

    Evidence and revision conditions

    Passage in the article

    What would change this assessment: An economic calculation using the actual lab’s hardware access, algorithmic progress and release incentives.

    E² editorial review · September 25, 2026 · Revision 1

  22. F22

    A predominantly domestic frontier-performance run is plausible within twelve months of publication.

    Type
    Forecast
    Support
    Conjectural
    Review
    Not yet resolvable

    The implied deadline is August 6, 2027. The article gives no fixed benchmark suite, comparator set, domestic-content threshold or qualifying training stage. Without these, later events can be fitted to the forecast after the fact.

    Evidence and revision conditions

    Passage in the article

    Depends on F08, F17, F19.

    What would change this assessment: Preregister resolution criteria and a probability, then score against the August 6, 2027 outcome.

    E² editorial review · September 25, 2026 · Revision 1

  23. F23

    Performance per dollar matters most when defining the frontier.

    Type
    Value judgment
    Support
    Not applicable
    Review
    Scope identified

    This selects an objective. It can be appropriate for a user or operator, but it is not a universal empirical result. Capability, latency, reliability and affordability can lead different users to different choices.

    Evidence and revision conditions

    Passage in the article

    What would change this assessment: State the decision-maker and their objective; compare decisions across plausible priorities.

    E² editorial review · September 25, 2026 · Revision 1

Review history: September 25, 2026 — first published ledger and arithmetic audit. Original article unchanged; corrections remain open.