Files
SystemSimulationApp/tests/manual/analyze_real_skip.py
T
ljz 7611f13208 修复循环信号与事件采样并接入 LSTP 接触定位,补充八路验证及复用实验
相较上一版 Jacobian 确定性复用更新,本次补齐事件边界一致性、结果两侧采样及接触事件定位;保留已有物性复用和组件力学公式。

- 统一 UD00 信号求值与下一事件查询的绝对时间边界,修复循环边界浮点舍入导致的阶段错位、重复或漏报,并覆盖零时长、多阶段及长周期场景。
- 引入原生输出语义 v2:保留规则网格真实时间,补充内部时间事件和状态事件的左邻及事件后采样,按保存时间、状态和离散模式重放结果。
- 两条代码生成路径均发出 LSTP 接触描述,默认定位间隙过零及非负力模式的力截断;仅在接受事件时更新防重复记录,增加 contactEvents 诊断计数。
- 补充 MASS/LSTP 独立事件实验、八路全曲线与驱动阶段配对评估,以及 Amesim 不连续点输出对照和力差定位报告;MASS 新增释放机制仍保留为独立实验。
- 保存局部 probe、context 访问与回退、shadow replay、R288 real skip/typed replay 及阀门数值尾部诊断工具和报告;未证明净收益的实验不启用为生产默认优化。
- 更新原生运行说明和元件建模规范,补充信号边界、输出语义、接触事件和实验依赖回归测试。

验证:五组专项回归共 34 项全部通过;37 个待提交 Python 文件语法检查通过;git diff --cached --check 通过。
2026-09-17 23:50:13 +08:00

145 lines
15 KiB
Python
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
"""Summarize independent serial measurements; audit times are never performance data."""
from pathlib import Path
import json,statistics
import real_skip_experiment as experiment
OUT=experiment.OUT
def load(path):return json.loads(path.read_text(encoding='utf-8'))
def write(path,value):experiment.write(path,value)
def compare_bytes(actual,reference):
offset=0
with actual.open('rb') as a,reference.open('rb') as b:
while True:
x=a.read(1024*1024);y=b.read(1024*1024)
if x!=y:
local=next((i for i,(u,v) in enumerate(zip(x,y)) if u!=v),min(len(x),len(y)))
different=offset+local
record={'file':str(actual),'reference':str(reference),'firstByte':different}
if actual.name=='jacobians.bin':
width=(1+132+132*132)*8;j,within=divmod(different,width);element=within//8
record.update(jacobian=j+1,field='t' if element==0 else 'y' if element<=132 else 'Jacobian',element=element)
if element>132:record.update(row=(element-133)%132,column=(element-133)//132)
write(OUT/'first-file-mismatch.json',record)
raise AssertionError(record)
if not x:break
offset+=len(x)
return offset
def main():
checks=[]
for label in ['audit-control-run','audit-skip-run','audit-forced-reject']:
for name in ['jacobians','states','outputs','events']:
ref=experiment.BASE.parent/('all-audit' if name=='jacobians' else 'all-run-0')/(name+'.bin')
n=compare_bytes(OUT/label/(name+'.bin'),ref)
checks.append(dict(run=label,field=name,bytes=n,exact=True))
write(OUT/'byte-comparison.json',checks)
rows=load(OUT/'performance.json');audit=load(OUT/'audit-skip-run/validation.json');forced=load(OUT/'audit-forced-reject/validation.json')
assert len(rows)==14 and all(r['exact'] for r in rows)
a=[r for r in rows if not r['skip']];b=[r for r in rows if r['skip']]
assert all(r['counters']['realSkips']==896 and r['counters']['nativeOriginalExecutions']==0 and r['counters']['rejects']==0 for r in b)
def timing(row,key):return row[key] if key in row else row['counters'][key]
keys=['originalSeconds','validationSeconds','overlayPatchSeconds','commitSeconds','fallbackSeconds','pathSeconds','baselineRecordSeconds','jacobianSeconds','solveCpuSeconds','solveSeconds']
def stat(values):return dict(median=statistics.median(values),min=min(values),max=max(values))
summary={side:{k:stat([timing(r,k) for r in rs]) for k in keys} for side,rs in [('control',a),('skip',b)]}
pairs=[]
for i in range(1,8):
c=next(r for r in a if r['label']==f'pair-{i}-control');s=next(r for r in b if r['label']==f'pair-{i}-skip')
stage=sum(s['counters'][k] for k in ['validationSeconds','overlayPatchSeconds','commitSeconds'])
pairs.append(dict(pair=i,originalUs=c['counters']['originalSeconds']/896*1e6,replayStagesUs=stage/896*1e6,
targetNetSavingSeconds=c['counters']['pathSeconds']-s['counters']['pathSeconds'],
originalMinusReplaySeconds=c['counters']['originalSeconds']-stage,
incrementalMetadataSeconds=s['counters']['baselineRecordSeconds']-c['counters']['baselineRecordSeconds'],
jacobianDeltaSeconds=s['jacobianSeconds']-c['jacobianSeconds'],
integrationCpuDeltaSeconds=s['solveCpuSeconds']-c['solveCpuSeconds'],integrationWallDeltaSeconds=s['solveSeconds']-c['solveSeconds']))
summary['paired']={k:stat([r[k] for r in pairs]) for k in pairs[0] if k!='pair'}
summary['pairs']=pairs;write(OUT/'performance-summary.json',summary)
med=lambda side,k:summary[side][k]['median']
orig=summary['paired']['originalUs']['median'];replay=summary['paired']['replayStagesUs']['median']
lines=['# R288 / position16 最小 real skip 实验','',
'## 结论','',
'1. **完整求解结果逐位一致。** 0–10 s 全轨迹,896 个 132×132 Jacobian(15,611,904 个元素)及每个 t/y、states、outputs、事件、最终状态与既有基线一致;没有使用数值容差。',
'2. **896/896 次真实 skip;自然 reject/fallback 为 0。** commit 896 次,原目标 native operation 实际执行 0 次,Reference 双路执行 0 次。',
f'3. **当前 replay 更贵。** 7 轮中位数:原 operation {orig:.3f} µs/次,validation + overlay/patch + commit {replay:.3f} µs/次({replay/orig:.2f} 倍)。目标完整路径 control {med("control","pathSeconds")/896*1e6:.3f} µs/次,real skip {med("skip","pathSeconds")/896*1e6:.3f} µs/次。',
'4. **暂不把这一实现直接扩展到 position52。** 正确性门槛已满足,净收益门槛未满足。应先降低 metadata 采集、事务复制和解释执行成本;本轮不能据此推断 position52 或其他 interval 的收益。','',
'## 范围与实现','',
'- 全部改动仅在独立生成的实验 worker 及 `tests/manual`;生产路径、默认开关、property cache 语义、原 whole-context guard 未改动。',
'- 入口仅为 `lp_color == 6 && region == 288` 的 `case 16`,位于原 `lp_reuse` 失败之后。其他位置继续原执行。',
'- baseline position16 用单独命名空间的 kernels 采集必需的有序事件 metadata。其他 operation 使用原 kernels,没有全局访问插桩。',
'- guard 来自已验证的 shadow 源码,构建时校验其 SHA-256,并断言 guard 函数保持一致;唯一变量替换是将原事务入口 count 改为当前真实 probe 入口 count。',
'- 复制当前 probe 到私有 overlay;按当前 entries 验证 first-match/miss,重定位逻辑 slot,从当前 count 追加。通过后构造地址已转换的有序 write set,再 commit。已有 entries、未写字段和其他 pipe slots 保持原值。',
'- Observer、capacity/scratch、memo lifetime/value、消费字段、valid、pipe branch、未知副作用/非有限值等原保护条件保留;失败只回退原 operation。',
'- 性能版不执行 Reference、不导出访问日志或 Jacobian、不逐次比较完整 context。保留运行所需 metadata、严格 guard、overlay/patch、计时和累计计数。','',
'## 正确性证据','',
'| 检查 | 结果 |','|---|---:|',
'| Jacobian、t/y 与既有 all-audit 文件逐字节比较 | 896/896 一致 |',
'| 每个 Jacobian 的 baseline + 27 group evaluator 出口 | 25,088/25,088 一致 |',
'| 目标 operation 入口 / 出口 | 896 / 896 一致 |',
'| 返回状态、dy/w、active property entries 有序字段/valid、全部 pipe slots | 逐位一致 |',
'| memo 每次 Jacobian 的表内容、owner 绑定、recording 生命周期 | 一致 |',
'| 所有 probe 的 memo entries 只读检查 | 通过 |',
'| gas memo 内容、kernel 绑定及计数 | 一致 |',
'| errno、x87/SSE flags/rounding、warning/observer | 一致 |',
'| count 差异 / slot relocation 的目标 probe | 638 / 571 |',
'| 顺序 append | 1,658 |',
'| PT miss → PT miss / 近零流量无查询 | 829 / 67 |','',
'Audit 对照来自**独立 control 进程真实执行当前 probe**的出口,约 1.88 GB 二进制数据。跨进程地址比较采用 owner/function 绑定身份,数值字段保持原始位模式;不把 C padding 当成数值。memo 表逐 Jacobian 比较,后续每个 probe 同时检查整表未变。目标入口、出口与所有 group 的完整 evaluator 出口均参与比较。audit 的 I/O 保留并恢复 errno、x87 和 SSE 环境。','',
'| 求解器计数 | control / real skip |','|---|---:|',
'| accepted / rejected | 10840 / 918 |','| Newton iterations / convergence failures | 19371 / 798 |',
'| nfev / njev / nlu | 44467 / 896 / 3106 |','| solverStarts / stateTransitions | 4 / 1 |','',
'额外负例:每 128 个目标 probe 注入一次晚期 nonfinite-output reject,发生在 overlay 已执行有序更新之后。7 次 reject 均完成无污染检查,fallback/native 原执行各 7 次,commit/skip 各 889 次;全轨迹仍与同一 control 和既有基线逐位一致。没有放宽 guard,也未改变容差。','',
'## 性能方法与原始数据','',
'Windows / MinGW GCC,原构建优化选项(`-O3 -ffp-contract=off -fno-fast-math`)。完成正确性和回退测试后单独编译性能 worker;control、skip 各预热一次,再交替串行运行 7 组。表中顺序就是实际顺序。每轮 states/outputs/events 指纹和全部指定求解器计数保持一致,skip 每轮均为 896、reject 0、原执行 0。','',
'QPC 粗粒度计时;累计整数 tick,积分过程中不做浮点时间换算。validation 包括语义校验所必需的临时有序更新;overlay/patch 包括当前 context 复制及提交 write set 构造;commit 是真实字段写回。完整 target 路径另计,包含调度、计数和额外计时开销。未减去计时器自身成本,未使用 shadow/audit 时间推断性能。','',
'| 运行顺序 | 原执行 µs/次 | validation µs/次 | overlay/patch µs/次 | commit µs/次 | fallback ms | target 总 ms | baseline pos16 总 ms | Jacobian s | 积分 CPU s | 积分 wall s | skip/reject |',
'|---|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|']
for r in rows:
c=r['counters'];us=lambda k:c[k]/896*1e6
lines.append(f'| {r["label"]} | {us("originalSeconds"):.3f} | {us("validationSeconds"):.3f} | {us("overlayPatchSeconds"):.3f} | {us("commitSeconds"):.3f} | {c["fallbackSeconds"]*1e3:.3f} | {c["pathSeconds"]*1e3:.3f} | {c["baselineRecordSeconds"]*1e3:.3f} | {r["jacobianSeconds"]:.6f} | {r["solveCpuSeconds"]:.6f} | {r["solveSeconds"]:.6f} | {c["realSkips"]}/{c["rejects"]} |')
lines+=['','baseline pos16:control 为原 baseline operation;skip 包含原 baseline operation + 本次 replay 必需 metadata 采集,不能漏算这部分成本。性能各轮没有自然 reject,因此 fallback 总时间为 0;这不代表一次 fallback 的成本为零,本轮未估计该分支的单次性能。','',
'### 中位数及 min/max','', '| 项目 | control 中位数 [min, max] | real skip 中位数 [min, max] |','|---|---:|---:|']
for k in keys:
def cell(side):
v=summary[side][k];return f'{v["median"]:.9f} [{v["min"]:.9f}, {v["max"]:.9f}]'
lines.append(f'| {k}(s,896 次累计) | {cell("control")} | {cell("skip")} |')
lines+=['','### 配对差值','',
'| 组 | 原计算 − replay 三阶段 ms | target 完整路径净节省 ms | 新增 metadata 采集 ms | Jacobian Δ s | 积分 CPU Δ s | 积分 wall Δ s |','|---|---:|---:|---:|---:|---:|---:|']
for r in pairs:lines.append(f'| {r["pair"]} | {r["originalMinusReplaySeconds"]*1e3:.3f} | {r["targetNetSavingSeconds"]*1e3:.3f} | {r["incrementalMetadataSeconds"]*1e3:.3f} | {r["jacobianDeltaSeconds"]:.6f} | {r["integrationCpuDeltaSeconds"]:.6f} | {r["integrationWallDeltaSeconds"]:.6f} |')
paired=summary['paired']
lines+=['','净节省为正表示节省,Δ = skip − control。','',
f'- 原计算成本 − replay 三阶段成本:配对中位数 **{paired["originalMinusReplaySeconds"]["median"]*1e3:.3f} ms / 896 次**。',
f'- target 完整路径净节省:配对中位数 **{paired["targetNetSavingSeconds"]["median"]*1e3:.3f} ms / 896 次**;范围 [{paired["targetNetSavingSeconds"]["min"]*1e3:.3f}, {paired["targetNetSavingSeconds"]["max"]*1e3:.3f}] ms。',
f'- 此外 baseline metadata 采集增加:配对中位数 **{paired["incrementalMetadataSeconds"]["median"]*1e3:.3f} ms**。',
f'- Jacobian callback 配对 Δ:中位数 {paired["jacobianDeltaSeconds"]["median"]:.6f} s,范围 [{paired["jacobianDeltaSeconds"]["min"]:.6f}, {paired["jacobianDeltaSeconds"]["max"]:.6f}] s。',
f'- 积分 wall 配对 Δ:中位数 {paired["integrationWallDeltaSeconds"]["median"]:.6f} s,范围 [{paired["integrationWallDeltaSeconds"]["min"]:.6f}, {paired["integrationWallDeltaSeconds"]["max"]:.6f}] s。',
'','全局时间受调频、调度和系统负载影响,不能把某轮 Jacobian/积分变快归因于这个 operation。目标路径在全部配对中均更慢,已经足以否定当前实现的局部净收益;本轮不预测整个 local probe 的最终加速比例。','',
'## 成本解释及下一阶段条件','',
'原 operation 在现有 Jacobian memo 的只读复用环境下已经很便宜;当前严格 replay 仍要初始化/复制 156,816 B 的 overlay(含完整 memo),解释事件并构造 write set。运行所需有效 metadata 为 2,224–23,344 B,Plan 固定预留 90,224 B,patch 预留 16,384 B。这些是当前隔离实现的实际成本,不是机制理论上的下限。',
'','因此,本轮证明了目标路径可以安全真实跳过,但没有证明性能优化成立。下一步应先针对事务存储和 metadata 表达做最小化,再重复同样的正确性与交替性能验收;不因 position16 正确就直接扩大到 position52/整个 R475/全部 reuse interval。','',
'## 复现与证据位置','',
'运行目录:`test/r288-real-skip-20260917/`。依赖上一轮生成的 local-probe worker、context-access worker 和 shadow-certified 源码;工具链复用原实验配置,不安装依赖。','',
'```powershell',
'.venv-win/Scripts/python.exe tests/manual/real_skip_experiment.py prepare --audit',
'.venv-win/Scripts/python.exe tests/manual/real_skip_experiment.py prepare --audit --skip',
'.venv-win/Scripts/python.exe tests/manual/real_skip_experiment.py run --audit',
'.venv-win/Scripts/python.exe tests/manual/real_skip_experiment.py run --audit --skip',
'.venv-win/Scripts/python.exe tests/manual/real_skip_experiment.py run --audit --skip --label audit-forced-reject --force 128',
'.venv-win/Scripts/python.exe tests/manual/real_skip_experiment.py prepare',
'.venv-win/Scripts/python.exe tests/manual/real_skip_experiment.py prepare --skip',
'.venv-win/Scripts/python.exe tests/manual/real_skip_experiment.py benchmark --pairs 7',
'.venv-win/Scripts/python.exe tests/manual/analyze_real_skip.py','```','',
'- `audit-{control,skip}-run/validation.json`:正确性、独立执行计数、完整求解器计数及文件指纹。',
'- `audit-forced-reject/validation.json`:真实回退及 rollback 检查。',
'- `byte-comparison.json`:全部 Jacobian/t/y、states、outputs、events 与原始基线的逐字节比较。',
'- `audit-control-run/audit.bin`:所有 group 和目标 operation 的实际执行对照出口。',
'- `performance.json`:14 次按实际顺序记录的原始数据;各轮目录保留独立日志和结果。',
'- `performance-summary.json`:中位数、min/max 和逐对差值。',
'- `{audit,perf}-{control,skip}/build.json`:源文件哈希、guard 一致性和构建记录。','']
report=Path(__file__).with_name('r288_real_skip_report.md');report.write_text('\n'.join(lines),encoding='utf-8')
print(json.dumps(summary['paired'],ensure_ascii=False,indent=2));print(report)
if __name__=='__main__':main()