定位不是求助,是通报——我们自己有防护、自己抓到了、自己定了性,其中还包括一个我们自己推翻的误判。下面这段英文可以直接复制粘贴发出去;后半页是中文对照和它为什么这么写。
来源:00-admin/MSG-ORGANISERS-2026-09-08-SERVO.md · 收件人:Demo Day 主办方 · 语气:平级通报,不是求援
Hi — a heads-up on a servo incident on our robot last night, and one small ask at the end.
WHAT HAPPENED
During an unattended remote teleoperation session, our left arm's shoulder-lift servo
reached 61 C. Our thermal guard stopped the session at that point (it warns at 50 C and
stops at 60 C, both well below the servo's own 70 C limit) — deliberately, so that we
park the arm rather than have the servo drop torque on its own and let the arm fall.
Root cause was not the servo. The arm had been driven into an extended pose and then
simply held there: 6821 of that session's 6889 control ticks (99%) were "holding
position". Nothing in our software made it stand down, so the joint carried the arm's
weight continuously until it overheated.
WHAT THE DATA SHOWS
Same-pose comparison, both arms parked at home, logged automatically:
left shoulder_lift 37 C / load 224 / 12.1 V
right shoulder_lift 36 C / load 216 / 12.0 V
So there is no hardware asymmetry between the two arms — 1 C apart at the same pose.
The overheating was caused by the pose and by nobody being there, not by a bad part.
Immediately after the 61 C event the servo did misbehave: driving it 2.6 degrees produced
15-33 bus communication failures, while its neighbours on the same bus, at HIGHER load,
had none. We initially read that as a failed servo. It was not. After roughly an hour with
torque disabled we ran a load-gradient sweep and it came back completely clean:
amplitude angle ID2 fails/load ID3 fails/load ID4 fails/load
5 0.4 deg 0 / 32 0 / 36 0 / 37
10 0.9 deg 0 / 78 0 / 77 0 / 81
15 1.3 deg 0 / 149 0 / 120 0 / 116
20 1.8 deg 0 / 195 0 / 148 0 / 143
25 2.2 deg 0 / 239 0 / 171 0 / 164
30 2.6 deg 0 / 255 0 / 207 0 / 188
40 3.5 deg 0 / 291 0 / 223 0 / 220
ID2 is the joint that overheated; ID3/ID4 are healthy references on the same bus, same
power, same wiring. ID2 now sustains a HIGHER load (291) than either reference (223/220)
with zero failures, including at the exact amplitude that failed 15-33 times while it was
still hot internally. Three further repeats at that amplitude: zero failures.
Conclusion: thermal derating, fully recovered. The servo is fine.
WHAT WE ALREADY HAD, AND WHAT WE ADDED
Already in place: a thermal guard that samples every servo round-robin and stops the
session at 60 C, with a re-read confirmation so one corrupted byte cannot stop a session.
Added last night:
- per-servo LOAD and VOLTAGE are now logged alongside temperature, at zero extra bus
traffic (the guard cycles through the three registers instead of only reading one).
This is what let us separate "this joint is working hard" from "this joint is broken";
- an idle auto-park: after 3 minutes with no operator input the arms retract slowly to
the home pose and hold there, keeping the gripper state so a held cup is not dropped,
and keeping torque on so the arm does not fall. Any keypress aborts it immediately.
Measured basis: at home the same joint sits at 37 C / load 224 and stays there.
THE ASK
Two things, neither urgent:
1. If anyone experienced with these SO-101 arms is around, could they spend ten minutes
checking the mechanical resistance of our left shoulder-lift joint by hand, with the
power off, against the right one? A servo was already replaced at that position once
before. Extra mechanical resistance would make that joint run hotter than its twin for
the same work, and no amount of software protection fixes that.
2. If a spare Feetech STS/SCS servo happens to be available to keep on the shelf, we
would appreciate knowing where to find one — not to fit now, but so a failure a few
days before Demo Day does not become a blocker.
Everything is safe right now: torque disabled on all 17 servos, all of them at 32-36 C.
三段式,顺序是刻意的:先说我们的防护起作用了,再说根因不是硬件,最后才提请求。
如果开头就写"我们的舵机坏了",读到后面主办方已经在想备件和风险;而事实是这颗舵机没坏,我们的保护按设计动作了,会话是被我们自己主动停下的,不是烧到掉力矩、手臂砸下去。这两件事在别人眼里的分量差很远。
其中包含一个我们自己推翻的结论,而且没有藏起来。热态时我判定这颗舵机劣化、该换,话都写好了。操作者一句「你应该做一个梯度测试才行呀」把它推翻——凉透之后它扛住的负载比两个健康邻居更高。消息里原样写了这个过程。
把误判写进去不是坦白癖:主办方拿到的每个数字,都会因为这一段而更可信。一封只报好消息的通报,读者没有办法判断哪些是量出来的、哪些是希望。
| 关节 | 温度 | 负载 | 电压 |
|---|---|---|---|
| left shoulder_lift | 37 °C | 224 | 12.1 V |
| right shoulder_lift | 36 °C | 216 | 12.0 V |
两条臂都停在 home、姿势相同时相差 1 °C。所以 61 °C 不是"这条臂天生差",是姿势 + 没人在。
| 幅度 | 角度 | ID2 失败/负载 | ID3 失败/负载 | ID4 失败/负载 |
|---|---|---|---|---|
| 5 | 0.4° | 0 / 32 | 0 / 36 | 0 / 37 |
| 10 | 0.9° | 0 / 78 | 0 / 77 | 0 / 81 |
| 15 | 1.3° | 0 / 149 | 0 / 120 | 0 / 116 |
| 20 | 1.8° | 0 / 195 | 0 / 148 | 0 / 143 |
| 25 | 2.2° | 0 / 239 | 0 / 171 | 0 / 164 |
| 30 | 2.6° | 0 / 255 | 0 / 207 | 0 / 188 |
| 40 | 3.5° | 0 / 291 | 0 / 223 | 0 / 220 |
ID2 就是过热那颗,ID3/ID4 是同总线、同供电、同线缆的健康对照。凉透后 ID2 的峰值负载 291 高于两个对照的 223 / 220,全程零失败——包括热态时失败 15–33 次的那个 2.6° 幅度。同幅度再重复三次,仍然零失败。
复现命令见 带载梯度扫描工具 那一页。
请人用手对比一下左右肩抬关节的机械阻力(断电,十分钟)。
这是唯一远程做不了的检查。同一位置已经换过一颗舵机——如果那里有额外机械阻力,同样的活它就是会比孪生关节更热,而这一点任何软件保护都修不了,只会一次次触发保护。所有能远程做的判别我都做完了,剩下的只有手感。
问一句备用 Feetech STS/SCS 舵机去哪找——不是现在要装。
提这个不是因为现在缺件(这颗好的)。是因为距离 Demo Day 只剩五天,而那几天没有人在现场:真到那时候坏了,"去哪买"这个问题本身就会变成卡点。现在问只要一句话。
数据来自 2026-09-08 凌晨的实测日志。原文 00-admin/MSG-ORGANISERS-2026-09-08-SERVO.md。
相关:过热事件全过程与被推翻的误判 · 带载梯度扫描工具 · 当晚修掉的三个遥操作 bug