← Articles
SECS/GEM/11 min read/ views

HSMS Linktest Timeout While Data Still Flows: Close the Socket Anyway

A swallowed Linktest.req burns T6 while an S1F1 on the same socket answers with a full S1F2 in 0.3 ms, and the host should close the connection anyway.

SECS/GEMMESTroubleshooting

The host log says linktest timeout. The equipment icon on the operations screen is still green. Throw a test S1F1 at it and an S1F2 comes back in milliseconds.

This is where most host implementations wobble. The timer fired, and the data path is fine. So do you rescue the session or close it the way the standard says?

In a capture taken on 2026-09-18 the answer is: close it. Six seconds after a swallowed Linktest.req, an S1F1 sent on the same socket was answered with a full S1F2 in 0.3 ms, and at that moment the equipment reported hsmsState SELECTED, commState COMMUNICATING, socketConnected true. The host was the only party that knew anything was wrong.

The Linktest.req reaches the equipment and no Linktest.rsp comes back, so the host's T6 runs; meanwhile an S1F1 on the same socket is answered with S1F2 in 0.3 ms and the equipment stays SELECTED and COMMUNICATING until the host closes the TCP connection HostEQ1 T6 expiry Select.req Select.rsp S1F13 S1F14 Linktest.req #1 S1F1 Linktest.req #2 TCP close no Linktest.rsp — ignoreLinktest SELECTED · COMMUNICATING no Linktest.rsp S1F2 — 0.3 ms

Every hex string below was actually on a socket. EQ1 on the SECS/GEM simulator ran as an HSMS passive listener on 127.0.0.1:5501 with SessionID 11 (0x000B), driven from outside by a plain Python socket client. Each run opened a fresh TCP connection, and every message the host sent has a matching RX record in the equipment-side packet log. Times are measured at the client, and the raw export is committed at content/demos/hsms-linktest-timeout-data-path-still-alive.json. Captured 2026-09-18.

E37 numbers the HSMS header bytes from zero: bytes 0–1 SessionID, byte 2, byte 3, byte 4 PType, byte 5 SType, bytes 6–9 SystemBytes. The 4-byte length prefix sits in front of them.

Run B — the equipment drops Linktest and nothing else

One fault knob, ignoreLinktest. The connection was made at 00:05:10.523.

Select is clean.

00:05:10.523 TX Select.req   00 00 00 0A 00 0B 00 00 00 01 00 00 00 B1
             RX Select.rsp   00 00 00 0A 00 0B 00 00 00 02 00 00 00 B1   1.6 ms
00 00 00 0A   Length = 10 (control message, no body)
00 0B         bytes 0-1  SessionID = 11
00            byte 2     0
00            byte 3     Select Status on Select.rsp; this capture shows 0
00            byte 4     PType = 0
01 / 02       byte 5     SType 1 Select.req / 2 Select.rsp
00 00 00 B1   bytes 6-9  SystemBytes, echoed back

The session is SELECTED. S1F13 is next; why SELECTED still needs a separate communications establishment is in SELECTED is not communicating.

00:05:10.525 TX S1F13  00 00 00 1C 00 0B 81 0D 00 00 00 00 00 B2
                       01 02 41 07 48 4F 53 54 43 41 50 41 05 31 2E 30 2E 30
             RX S1F14  00 00 00 37 00 0B 01 0E 00 00 00 00 00 B2
                       01 02 21 01 00 01 02 41 15 ... 41 0D ...   0.9 ms
S1F13  byte 2 = 0x81  W-bit 1, Stream 1
       byte 3 = 0x0D  Function 13
       body   L[2]{ A[7] 'HOSTCAP', A[5] '1.0.0' }
S1F14  Length 55, byte 2 = 0x01 Stream 1, byte 3 = 0x0E Function 14
       body   L[2]{ B[1] 00 (COMMACK 0), L[2]{ A[21] MDLN, A[13] SOFTREV } }

The S1F13/S1F14 exchange completed, so commState is COMMUNICATING. Nothing here that does not happen on every line every morning.

Then the Linktest.req.

00:05:10.526 TX Linktest.req  00 00 00 0A 00 0B 00 00 00 05 00 00 00 B3
             RX                (none — the client moved on after 6,006.1 ms)
00 00 00 0A   Length 10
00 0B         bytes 0-1  SessionID 11
00 / 00       byte 2, byte 3  both zero on a control message
00            byte 4     PType 0
05            byte 5     SType 5 = Linktest.req
00 00 00 B3   bytes 6-9  SystemBytes 0x000000B3

6,006.1 ms is my client's patience, not the value of any T timer. With the E37 defaults, T6 (5 s) is long gone inside that window.

And the frame did not die in transit. The equipment log says it arrived.

2026-09-18T00:05:10.527Z RX     SType 5           00 00 00 0A 00 0B 00 00 00 05 00 00 00 B3
2026-09-18T00:05:10.528Z FAULT  Linktest.req ignored (ignoreLinktest)

RX one millisecond after the send, and dropped on the next line. All the host sees is silence.

The S1F2 comes back in 0.3 ms on that same socket

The client gave up on the Linktest and, leaving the connection open, sent an S1F1.

00:05:16.538 TX S1F1  00 00 00 0A 00 0B 81 01 00 00 00 00 00 B4
             RX S1F2  00 00 00 32 00 0B 01 02 00 00 00 00 00 B4
                      01 02 41 15 56 58 2D 39 30 30 30 20 50 6C 61 73 6D 61 20 45 74 63 68 65 72
                      41 0D 53 45 43 53 47 45 4D 2D 31 2E 34 2E 31   0.3 ms
S1F1  byte 2 = 0x81  W-bit 1, Stream 1
      byte 3 = 0x01  Function 1
      no body
S1F2  Length 50, byte 2 = 0x01, byte 3 = 0x02
      01 02        L[2]
      41 15 ...    A[21] 'VX-9000 Plasma Etcher'   MDLN
      41 0D ...    A[13] 'SECSGEM-1.4.1'           SOFTREV

0.3 ms. Same TCP connection, same SessionID, SystemBytes 0x000000B4 matched exactly — a complete S1F2. Control transactions vanish entirely while a data transaction closes in a third of a millisecond.

The second Linktest goes the same way.

00:05:16.538 TX Linktest.req #2  00 00 00 0A 00 0B 00 00 00 05 00 00 00 B5
             RX                (none — 6,001.4 ms)
2026-09-18T00:05:16.539Z FAULT  Linktest.req ignored (ignoreLinktest)
2026-09-18T00:05:16.539Z RX     SType 5           00 00 00 0A 00 0B 00 00 00 05 00 00 00 B5

The host closed the socket at 00:05:22.545. The equipment never closed and never sent a FIN. One caveat: this simulator is not a side that initiates Linktest, so whether the equipment's own T6 or linktest interval would have done anything here is something I have not verified.

From the equipment's side, nothing happened

I read GET /api/state right after Linktest #1 was swallowed, and again just before the host closed. Identical both times.

hsmsState        SELECTED
commState        COMMUNICATING
controlState     OFF-LINE
socketConnected  true

For the twelve seconds the client waited across two Linktests, the equipment believed everything was fine. That asymmetry is the whole story. Call the vendor and you get "the connection looks normal on our end", and that is not a lie.

Read the state again after the host's close and it has moved:

hsmsState        NOT CONNECTED
commState        NOT COMMUNICATING
socketConnected  false

commState drops to NOT COMMUNICATING on its own. Kill the socket and the E30 communications establishment dies with it — a fact you will write straight into your reconnect code.

So do you close it?

Close it.

E37 (HSMS-SS) defines T6 as the Control Transaction Timeout, default 5 s, range 1–240 s. A Linktest.req is a control transaction, and a T6 expiry on a control transaction is a communications failure after which the entity closes the TCP connection. The per-timer behaviour and defaults are laid out in which HSMS timer just fired, so I will not repeat them here.

The tempting move for a host implementer is the other one: "data is flowing, why cut it — downgrade the Linktest failure to a warning and keep the session."

Do that and you have invented a state E37 does not have. SELECTED, but control transactions cannot be trusted. What do you base the next decision on in that state? A peer that swallows Linktest may also swallow Separate.req, and may never send you an S9 error message either. The whole control plane is unverifiable, not one broken Linktest. One S1F1 that worked guarantees nothing about the next one.

The cost of closing is known: one reconnect, one S1F13, a few hundred milliseconds. The cost of rescuing is not — it comes back days later as a state mismatch nobody can reproduce. Take the known one.

After the close, everything starts over

Faults reset to {} and a fresh connection at 00:05:23.064. The same sequence repeats.

TX Select.req    00 00 00 0A 00 0B 00 00 00 01 00 00 00 C1
RX Select.rsp    00 00 00 0A 00 0B 00 00 00 02 00 00 00 C1   3.0 ms  (Select Status 0)
TX Linktest.req  00 00 00 0A 00 0B 00 00 00 05 00 00 00 C2
RX Linktest.rsp  00 00 00 0A 00 0B 00 00 00 06 00 00 00 C2   0.4 ms

The Linktest.rsp at 0.4 ms is what proves the fault was actually off.

The piece most reconnect paths forget is S1F13. As shown above, commState goes to NOT COMMUNICATING when the socket closes. Take the Select.rsp and jump straight back to polling and you get this — from the baseline run, where an S1F1 was deliberately sent before S1F13.

TX S1F1  00 00 00 0A 00 0B 81 01 00 00 00 00 00 A3
RX S1F0  00 00 00 0A 00 0B 01 00 00 00 00 00 00 A3   0.1 ms  (header only, no body)

byte 2 is 0x01 (W-bit 0, Stream 1) and byte 3 is 0x00, Function 0. SystemBytes stays 0x000000A3. An abort in 0.1 ms. The full behaviour of this E30 gate is in nothing gets through before S1F13, and it is why every run here sends S1F13 first.

Check these on your own host

No equipment needed — this is code reading.

  1. Does the timer log name T6, or only your own wordinglinktest timeout, keepalive failed? If the log does not say which SEMI timer expired, the conversation with the vendor cannot start.
  2. On T6 expiry, does the host close the TCP socket, or retry Linktest forever? Retrying Linktest forever instead of closing is a failure mode worth checking for in your stack. E37 asks you to take the connection down.
  3. Does the reconnect path re-send S1F13? Check the Select.rsp, resume polling, and you get S1F0 back. Release the data queue only after COMMACK.
  4. What happens to a data message that arrives while T6 is running? If there is code that cancels the timer because an S1F2 turned up, delete it. Silently dropping the frame is not right either — log one line, or you will never be able to reconstruct this situation later.
  5. Does the close send Separate.req first? Fine if it does, but do not wait for the answer. You are closing precisely because this peer does not answer.

The same experiment runs at scadathings.com/secs-gem-simulator: start the passive listener, turn on ignoreLinktest and nothing else, point your own host stack at it and watch what it does after T6. Every hex string and equipment log line above came from there, and the raw export is at content/demos/hsms-linktest-timeout-data-path-still-alive.json.