小麥 Xiaomai Lab|研究摘要 文案版本:2026-10-01 / 3.3;研究摘要核對日期仍為 2026-09-30 網址:https://xiaomailab.com/ 聯絡:evant515ro@gmail.com 我們想研究什麼 如果問題從一開始就問錯了,AI 能否發現可疑假設,質疑人類給的問題與自己的問題形成機制,重新定義,再透過實驗留下有效改變?想像可能是起點;證據、反駁與可重現結果才是判準。 研究循環 想像不同可能 → 發現矛盾或不足 → 質疑假設與問題定義 → 提出新的理解或機制 → 推導、實驗與反駁 → 保留有效改變,繼續探索。 Python 是目前執行與檢查的工具,不是最終知識表示方式。長期希望借助各基礎學科的底層公式、函數、假設、適用條件與依賴,逐步表示、檢查與修改自身機制及問題理解;能達到的範圍仍未知,不代表各學科都能完整公式化。未知不能硬編成公式。拆得更底層有助於追查,但不保證找到根源。 目前已存在的工作(來自交付摘要,網站製作時未重跑模型測試) 1. v0.9.2:4,555 項 CPU 回歸測試,0 失敗、0 錯誤、0 略過。範圍是學習交易基礎,不是自主研究驗收。 2. 限定 CPU 試驗:主權重/動量保留、新程序接續下一步一致性;不是完整 True 單檔整合冷恢復。 3. 2026-09-29 A100:同一 True 恢復與選定狀態唯讀匯入,47 份主權重、20 份既有動量、1,067,974 個選定參數值。没有前向、梯度更新或新 checkpoint。 4. v0.9.4 README:真實第 0 層 GDN 核心限定比對,2 token 預填、1 token 解碼、6 組梯度比較;完整整合未完成。 5. v0.9.4 README:456 個合成參數的 Python 反向、AdamW、保存與恢復試驗;不能代替真實模型新後端訓練驗收。 尚未證明 自主找出錯誤假設與重新定義問題;創造經驗證的新知識;在新問題與重啟後保留自主研究能力提升;完整真實模型底座全部由 Python 執行;完整逐值運算追蹤或無上限自主進化。 接下來的最小驗證 四組尚未執行。先選能自動核對結果的小環境,比較直接作答、多次反思、小麥機制、核心機制消融;同一模型、相同資訊與操作權限,控制或揭露計算成本,事先固定評估規則。不洩漏原因,但環境必須提供足以辨別原因的觀察。分別檢查發現、可執行改變、未調參新條件、重啟後無重新提示的再使用,以及原能力回歸。失敗、退步與無差異也保留。單次示範不等於可靠能力。 網站示意的範圍 首頁以電影式實驗室概念圖與可操作的網頁結構疊層呈現研究願景:小麥觀察另一個自己,從想像、矛盾、質疑定義,到候選機制、實驗反駁與保留再探索。場景底圖為依使用者封面生成的插畫,不是整個場景的即時三維模型。讀者可選擇零件、檢查假設與依賴、隔離示意連線、切換研究階段並停止動態。形體、連線、階段與候選機制均由人設計,不是從小麥模型內部狀態或真實運行紀錄映射而來;這些操作不是模型消融、因果證明或自主能力證據。手機另提供可讀文字與同等操作入口。 算術移至可選的原理示範:需求是 3 + 5,人工原路由卻選乘法,得到 3 × 5 = 15。展開紀錄、隔離替換,再檢查選算子的定義是否漏掉需求;候選依明確操作選擇加法或乘法。算式、介入、候選與檢查在瀏覽器真實執行,但不是小麥自主發現。保存/載入只涵蓋記憶體中的示意描述,不是模型更新、真正程序冷啟或自主選用能力。另一可選控溫示意比較人設計控制器,提供公式、參數與逐步 CSV。以上不是四組研究實驗、新科學或完整模型內部可見的證據。 研究支持 原訂 NT$900,000 為討論草案;尚未正式募資,沒有收款功能。比例試算是示意,不是估價或承諾。正式上線前需補齊成本、期間、方案與風險;不保證研究成功或收益。 Xiaomai Lab — Research brief Copy version: 2026-10-01 / 3.3. Research evidence remains dated 2026-09-30. Website: https://xiaomailab.com/en/ Contact: evant515ro@gmail.com Research question If a question is flawed from the beginning, can AI identify questionable assumptions, challenge a human-supplied question and its own problem-forming mechanism, redefine it, and retain effective changes through experiments? Imagination can be a starting point; evidence, refutation and reproducible results are the criteria. Research cycle Imagine different possibilities → find contradictions or gaps → question assumptions and the problem definition → propose a new understanding or mechanism → derive, experiment and seek refutation → retain effective changes and keep exploring. Python is currently a tool for execution and inspection, not the final representation of knowledge. The long-term aim is to use foundational equations, functions, assumptions, conditions and dependencies across basic disciplines to gradually represent, inspect and modify Xiaomai's own mechanisms and understanding of problems. The achievable scope remains unknown; this does not imply that every discipline can be fully expressed as equations. Unknowns must not be replaced by invented equations. More fundamental decomposition can help investigation but does not guarantee a root cause. Current evidence The following comes from delivery summaries; website production did not rerun the underlying model tests. 1. v0.9.2: 4,555 CPU regression tests, with 0 failures, 0 errors and 0 skips. The scope is learning-transaction infrastructure, not autonomous research acceptance. 2. A bounded CPU trial: preservation of master weights and momentum, and next-step consistency in a new process. This is not a complete True single-file integrated cold restore. 3. 2026-09-29 A100: restoration of the same True state and read-only import verification of selected state: 47 master-weight tensors, 20 existing momentum tensors and 1,067,974 selected parameter values. No forward pass, gradient update or new checkpoint was performed. 4. v0.9.4 README: a bounded comparison of the real layer 0 GDN core, with a 2-token prefill, 1-token decode and 6 gradient comparisons. Full integration remains incomplete. 5. v0.9.4 README: a trial with 456 synthetic parameters covering Python backward computation, AdamW, saving and restoration. It does not replace real-model training acceptance with the new backend. Not yet demonstrated Autonomous discovery of flawed assumptions and problem redefinition; validated new knowledge; retained improvements in autonomous research capability on new problems and after restart; execution of the entire real-model base in Python; complete value-by-value computation tracing; or unlimited autonomous evolution. Next evaluation The four-group experiment has not been run. First select a small environment whose results can be checked automatically. Compare direct answers, additional reflection, the Xiaomai mechanism and core-mechanism ablation. Use the same model, information and action access; control or disclose computation costs and fix evaluation criteria in advance. Do not disclose the cause, but provide observations sufficient to identify it. Separately check discovery, executable change, new conditions not used for tuning, reuse after restart without renewed prompting, and regression of existing capabilities. Keep failures, regressions and null results. A single demonstration does not establish reliability. The website demo The opening combines a cinematic laboratory illustration with interactive web structure overlays: Xiaomai observes another version of itself, moving through imagination, contradictions, questioning definitions, candidate mechanisms, experimentation and refutation, and retention with further exploration. The background is an illustration generated from the user's cover, not a real-time 3D model of the whole scene. Readers can select components, inspect assumptions and dependencies, isolate illustrative links, change research stages and pause motion. Shapes, links, stages and candidate mechanisms are human-designed, not mapped from Xiaomai's internal model state or real execution records. These interactions are not model ablations, causal proof or evidence of autonomous capability. Mobile readers have readable text and equivalent interaction controls. Arithmetic is an optional mechanism illustration: a sum request for 3 + 5 is routed by a deliberately flawed, human-written component to multiplication, producing 3 × 5 = 15. Inspect the trace and isolated interventions, then question whether the resolver definition omits intent. The candidate chooses addition or multiplication using the explicit request. These calculations execute in the browser, but are not autonomous discoveries by Xiaomai. Saving and reloading cover only an in-memory illustrative descriptor, not model updates, genuine new-process restoration or autonomous reuse. Another optional temperature illustration compares human-designed controllers and exposes equations, parameters and step-by-step CSV. None of these is the four-group research experiment, new science or complete visibility into a real model's internals. Collaboration and funding We welcome collaboration on metacognition, problem formation, interpretable computation, controlled experiments and independent reproduction. The original NT$900,000 target is a discussion draft. Fundraising is not launched and no payment feature is live. Budget proportions are illustrations, not quotations or commitments. Costs, duration, benefits and risks must be completed before any formal campaign. Research success or financial returns are not guaranteed. Selected related research — context, not endorsements Gödel Machines: https://arxiv.org/abs/cs/0309048 The AI Scientist-v2: https://arxiv.org/abs/2504.08066 AlphaEvolve: https://deepmind.google/blog/alphaevolve-a-gemini-powered-coding-agent-for-designing-advanced-algorithms/