HSMS Byte 4 Is PType, and a One-Byte Shift Turns Linktest Into Silence
A socket capture: the listener answers a wrong PType as if nothing happened, and a Linktest.req written one byte left gets no reply at all.
You send a Linktest.req and no Linktest.rsp comes back. T6 fires and the host closes a perfectly good socket. Then you open the equipment-side log and the frame is sitting right there. It arrived. The equipment just did not answer it.
An HSMS message is a 4-byte length prefix followed by a 10-byte header, and SEMI E37 numbers those header bytes from zero. Bytes 0-1 are the SessionID, bytes 2 and 3 are Header Byte 2 and 3, byte 4 is PType, byte 5 is SType, bytes 6-9 are the SystemBytes. PType says how the body that follows should be read, and in E37 the value 0 means SECS-II encoding. I have never seen anything but 0 in the field. SType picks the control message — Select.req, Linktest.req and so on — and 0 means this is a data message. The two bytes are adjacent. Write SType at header offset 4 instead of 5 and a Linktest.req puts its 5 in the PType slot, leaving 00 where SType belongs.
A wrong PType went straight through
Everything below was captured by opening a socket against a passive HSMS listener and acting as the host. The equipment side logged every frame as RX, so these are bytes that were really on the wire.
After Select and an S1F13 to open communication, I sent S1F1 twice down the same socket. Once with PType 0, once with PType 3.
TX S1F1 PType=0 00 00 00 0A 00 0B 81 01 00 00 00 00 10 03
RX S1F2 00 00 00 32 00 0B 01 02 00 00 00 00 10 03
01 02 41 15 56 58 2D 39 30 30 30 20 50 6C 61 73 6D 61 20 45 74 63 68 65 72
41 0D 53 45 43 53 47 45 4D 2D 31 2E 34 2E 31
An ordinary transaction. The request is header-only, 14 bytes; byte 2 is 81, so W-bit 1 and Stream 1, and byte 3 is 01 for Function 1. The reply carries an L[2] holding MDLN VX-9000 Plasma Etcher and SOFTREV SECSGEM-1.4.1.
TX S1F1 PType=3 00 00 00 0A 00 0B 81 01 03 00 00 00 10 04
RX S1F2 00 00 00 32 00 0B 01 02 00 00 00 00 10 04
01 02 41 15 56 58 2D 39 30 30 30 20 50 6C 61 73 6D 61 20 45 74 63 68 65 72
41 0D 53 45 43 53 47 45 4D 2D 31 2E 34 2E 31
The two requests differ in byte 4 — 00 against 03 — and in the last digit of the SystemBytes. The reply bodies are byte-for-byte identical. The listener never reads PType at all: its source branches on SType, decodes PType and then does nothing with it. And it stamps PType 0 on everything it sends, which is why offset 4 is 00 in both replies above.
The point is not that the listener is wrong. A field your test partner does not check is a field your tests cannot catch. If your encoder is writing rubbish into byte 4, the simulator run is green end to end. E37 defines a Reject.req (SType 7) as the answer to an unsupported PType, but I did not verify its reason-code numbering in this run, and I have no evidence about whether a given vendor stack performs the check. One thing is certain: if it does, the day you meet that equipment is your first day in the fab.
A shifted Linktest becomes S0F0
On the same socket, the next frame put SType 5 at offset 4 instead of offset 5.
TX shifted Linktest 00 00 00 0A 00 0B 00 00 05 00 00 00 10 05
RX (nothing) waited 3 s, nothing came back
The equipment logged this one as RX and parsed it as ptype: 5, stype: 0. The length prefix was right and the framing never broke. The receiver followed the rules. SType 0 means data message; byte 2 is 00, so W-bit 0 and Stream 0; byte 3 is 00, so Function 0. That is an S0F0 that does not ask for a reply. The listener only builds a reply when the W-bit is set. It had no reason to answer, so it did not.
The very next frame was the same Linktest written correctly.
TX Linktest.req 00 00 00 0A 00 0B 00 00 00 05 00 00 10 06
RX Linktest.rsp 00 00 00 0A 00 0B 00 00 00 06 00 00 10 06
Instant. The only difference between the two frames is whether 05 sits at offset 4 or offset 5. Same socket, same session.
What makes this nasty to debug is that the symptom reads as "the equipment does not answer Linktest". Start from that sentence and you go digging through equipment settings, firewalls and the vendor's T6 value. The culprit is your own header builder. And Select.req and Separate.req disappear the same way — whatever the SType, once it shifts it becomes a W-bit-less S0F0.
The whole capture
RX Select.req 00 00 00 0A 00 0B 00 00 00 01 00 00 10 01
TX Select.rsp 00 00 00 0A 00 0B 00 00 00 02 00 00 10 01
RX S1F13 W 00 00 00 0C 00 0B 81 0D 00 00 00 00 10 02 01 00
TX S1F14 00 00 00 37 00 0B 01 0E 00 00 00 00 10 02 01 02 21 01 00 ...
RX S1F1 PType=0 00 00 00 0A 00 0B 81 01 00 00 00 00 10 03
TX S1F2 00 00 00 32 00 0B 01 02 00 00 00 00 10 03 ...
RX S1F1 PType=3 00 00 00 0A 00 0B 81 01 03 00 00 00 10 04
TX S1F2 00 00 00 32 00 0B 01 02 00 00 00 00 10 04 ...
RX shifted Linktest 00 00 00 0A 00 0B 00 00 05 00 00 00 10 05
RX Linktest.req 00 00 00 0A 00 0B 00 00 00 05 00 00 10 06
TX Linktest.rsp 00 00 00 0A 00 0B 00 00 00 06 00 00 10 06
Direction is from the equipment's point of view: RX came up from the host, TX went out from the equipment. The missing TX line after the shifted Linktest is the entire article.
What to put in the host
Make PType a constant 0 in the encoder, not a variable. If there is no code path that writes a value into byte 4, there is nothing to shift.
In the decoder, do the opposite: read PType and drop the message if it is not 0. Parse it as SECS-II regardless and you will occasionally succeed, and invent a message nobody sent.
Then one unit test on the header builder. Build a Linktest.req and assert the byte string is 00 00 00 0A ss ss 00 00 00 05 .... A test that compares against fixed offsets breaks the moment this bug lands. The integration test will not break — as shown above, the other side is not looking at PType.
Last thing: when you log a "no reply", log the bytes you sent alongside it. For this failure that one hex line ends the investigation in three seconds.
The listener is up right now. Push the hex above into a socket and you get the same result: SECS/GEM simulator.