← reports index
对外沟通 · 2026-09-08

给主办方的通报:舵机过热,以及两个不紧急的请求

定位不是求助,是通报——我们自己有防护、自己抓到了、自己定了性,其中还包括一个我们自己推翻的误判。下面这段英文可以直接复制粘贴发出去;后半页是中文对照和它为什么这么写。

直接复制这段

来源:00-admin/MSG-ORGANISERS-2026-09-08-SERVO.md · 收件人:Demo Day 主办方 · 语气:平级通报,不是求援

Hi — a heads-up on a servo incident on our robot last night, and one small ask at the end.

WHAT HAPPENED
During an unattended remote teleoperation session, our left arm's shoulder-lift servo
reached 61 C. Our thermal guard stopped the session at that point (it warns at 50 C and
stops at 60 C, both well below the servo's own 70 C limit) — deliberately, so that we
park the arm rather than have the servo drop torque on its own and let the arm fall.

Root cause was not the servo. The arm had been driven into an extended pose and then
simply held there: 6821 of that session's 6889 control ticks (99%) were "holding
position". Nothing in our software made it stand down, so the joint carried the arm's
weight continuously until it overheated.

WHAT THE DATA SHOWS
Same-pose comparison, both arms parked at home, logged automatically:

    left  shoulder_lift  37 C / load 224 / 12.1 V
    right shoulder_lift  36 C / load 216 / 12.0 V

So there is no hardware asymmetry between the two arms — 1 C apart at the same pose.
The overheating was caused by the pose and by nobody being there, not by a bad part.

Immediately after the 61 C event the servo did misbehave: driving it 2.6 degrees produced
15-33 bus communication failures, while its neighbours on the same bus, at HIGHER load,
had none. We initially read that as a failed servo. It was not. After roughly an hour with
torque disabled we ran a load-gradient sweep and it came back completely clean:

    amplitude   angle    ID2 fails/load   ID3 fails/load   ID4 fails/load
        5       0.4 deg     0 /  32          0 /  36          0 /  37
       10       0.9 deg     0 /  78          0 /  77          0 /  81
       15       1.3 deg     0 / 149          0 / 120          0 / 116
       20       1.8 deg     0 / 195          0 / 148          0 / 143
       25       2.2 deg     0 / 239          0 / 171          0 / 164
       30       2.6 deg     0 / 255          0 / 207          0 / 188
       40       3.5 deg     0 / 291          0 / 223          0 / 220

ID2 is the joint that overheated; ID3/ID4 are healthy references on the same bus, same
power, same wiring. ID2 now sustains a HIGHER load (291) than either reference (223/220)
with zero failures, including at the exact amplitude that failed 15-33 times while it was
still hot internally. Three further repeats at that amplitude: zero failures.

Conclusion: thermal derating, fully recovered. The servo is fine.

WHAT WE ALREADY HAD, AND WHAT WE ADDED
Already in place: a thermal guard that samples every servo round-robin and stops the
session at 60 C, with a re-read confirmation so one corrupted byte cannot stop a session.

Added last night:
  - per-servo LOAD and VOLTAGE are now logged alongside temperature, at zero extra bus
    traffic (the guard cycles through the three registers instead of only reading one).
    This is what let us separate "this joint is working hard" from "this joint is broken";
  - an idle auto-park: after 3 minutes with no operator input the arms retract slowly to
    the home pose and hold there, keeping the gripper state so a held cup is not dropped,
    and keeping torque on so the arm does not fall. Any keypress aborts it immediately.
    Measured basis: at home the same joint sits at 37 C / load 224 and stays there.

THE ASK
Two things, neither urgent:

1. If anyone experienced with these SO-101 arms is around, could they spend ten minutes
   checking the mechanical resistance of our left shoulder-lift joint by hand, with the
   power off, against the right one? A servo was already replaced at that position once
   before. Extra mechanical resistance would make that joint run hotter than its twin for
   the same work, and no amount of software protection fixes that.

2. If a spare Feetech STS/SCS servo happens to be available to keep on the shelf, we
   would appreciate knowing where to find one — not to fit now, but so a failure a few
   days before Demo Day does not become a blocker.

Everything is safe right now: torque disabled on all 17 servos, all of them at 32-36 C.

为什么这么写

三段式,顺序是刻意的:先说我们的防护起作用了,再说根因不是硬件,最后才提请求。

如果开头就写"我们的舵机坏了",读到后面主办方已经在想备件和风险;而事实是这颗舵机没坏,我们的保护按设计动作了,会话是被我们自己主动停下的,不是烧到掉力矩、手臂砸下去。这两件事在别人眼里的分量差很远。

其中包含一个我们自己推翻的结论,而且没有藏起来。热态时我判定这颗舵机劣化、该换,话都写好了。操作者一句「你应该做一个梯度测试才行呀」把它推翻——凉透之后它扛住的负载比两个健康邻居更高。消息里原样写了这个过程。

把误判写进去不是坦白癖:主办方拿到的每个数字,都会因为这一段而更可信。一封只报好消息的通报,读者没有办法判断哪些是量出来的、哪些是希望。

数据(消息里那两张表的出处)

同姿势对照 —— 证明两条臂硬件上没有差别

关节温度负载电压
left shoulder_lift37 °C22412.1 V
right shoulder_lift36 °C21612.0 V

两条臂都停在 home、姿势相同时相差 1 °C。所以 61 °C 不是"这条臂天生差",是姿势 + 没人在。

梯度扫描 —— 推翻"它坏了"的那张表

幅度角度ID2 失败/负载ID3 失败/负载ID4 失败/负载
50.4°0 / 320 / 360 / 37
100.9°0 / 780 / 770 / 81
151.3°0 / 1490 / 1200 / 116
201.8°0 / 1950 / 1480 / 143
252.2°0 / 2390 / 1710 / 164
302.6°0 / 2550 / 2070 / 188
403.5°0 / 2910 / 2230 / 220

ID2 就是过热那颗,ID3/ID4 是同总线、同供电、同线缆的健康对照。凉透后 ID2 的峰值负载 291 高于两个对照的 223 / 220,全程零失败——包括热态时失败 15–33 次的那个 2.6° 幅度。同幅度再重复三次,仍然零失败。

复现命令见 带载梯度扫描工具 那一页。

两个请求,为什么是这两个

1

请人用手对比一下左右肩抬关节的机械阻力(断电,十分钟)。

这是唯一远程做不了的检查。同一位置已经换过一颗舵机——如果那里有额外机械阻力,同样的活它就是会比孪生关节更热,而这一点任何软件保护都修不了,只会一次次触发保护。所有能远程做的判别我都做完了,剩下的只有手感。

2

问一句备用 Feetech STS/SCS 舵机去哪找——不是现在要装。

提这个不是因为现在缺件(这颗好的)。是因为距离 Demo Day 只剩五天,而那几天没有人在现场:真到那时候坏了,"去哪买"这个问题本身就会变成卡点。现在问只要一句话。

中文对照

通报的三件事

  1. 左臂肩抬舵机昨夜到 61 °C,我们自己的温度守卫在限值处主动停机(警告 50 / 停机 60,都低于舵机自身的 70)
  2. 根因不是舵机:同姿势 37 °C/224 对 36 °C/216,差 1 °C。那次会话 6889 拍里 6821 拍(99%)是 holding position——手臂被开到伸展姿势后就一直撑着,软件里没有任何东西让它退出来
  3. 一度误判为损坏,被梯度扫描推翻

展示的能力

  • 温度守卫:轮询采样 + 二次确认,一个坏字节不会误停整场
  • 新增:负载和电压也进日志,而且总线流量一点没增加(三个寄存器轮换,不是多读两次)
  • 新增:空闲自动停放——3 分钟无操作则缓慢缩回 home,保持夹爪状态(手里有杯子不会扔)、保持力矩(卸了手臂会掉),按键立刻中止

这一页答不了什么

数据来自 2026-09-08 凌晨的实测日志。原文 00-admin/MSG-ORGANISERS-2026-09-08-SERVO.md。

相关:过热事件全过程与被推翻的误判 · 带载梯度扫描工具 · 当晚修掉的三个遥操作 bug