← Articles
SECS/GEM/10 min read/— views

Four SECS Requests in One write(): Pair Replies on SystemBytes, Never on Arrival Order

Four HSMS requests in one write() came back as three replies. Real captures of the missing W=0 reply, an S7F0 mid-batch, and the T3 that fires.

SECS/GEMMESTroubleshooting

Speed up a host polling loop by firing requests together and sooner or later a T3 timeout shows up one line at a time — while the equipment log shows every reply went out. In the capture below I sent four requests and got three replies. The missing one was never lost. One request in the batch was never going to be answered, and code that paired replies by arrival order slipped one slot from there on. The last request waits forever.

The host sends S1F3, S1F1, S1F11 and S1F3 in one write and the equipment returns only three replies — SystemBytes 0x4011, the S1F1 sent with W=0, gets no reply HostEquipment one write() S1F3 W=1 · 0x4010 S1F1 W=0 · 0x4011 S1F11 W=1 · 0x4012 S1F3 W=1 · 0x4013 S1F4 · 0x4010 S1F12 · 0x4012 S1F4 · 0x4013 0x4011no reply

Every capture below came off a socket I opened directly against the passive listener of the SECS/GEM simulator on 127.0.0.1:5501 — a real socket, not the simulator's own API. That is why the equipment side logged RX for each one, and that RX is the evidence these bytes actually crossed the wire. Each session starts with Select.req and S1F13 to reach COMMUNICATING.

Four requests at once, four replies

The healthy shape first. Four requests, W-bit set on all of them, 74 bytes in a single sendall().

TX (one write, 74 bytes)
00 00 00 12 00 0B 81 03 00 00 00 00 30 10 01 01 B1 04 00 00 00 02   S1F3  W=1  SystemBytes 0x3010
00 00 00 0A 00 0B 81 01 00 00 00 00 30 11                           S1F1  W=1  SystemBytes 0x3011
00 00 00 0C 00 0B 82 1D 00 00 00 00 30 12 01 00                     S2F29 W=1  SystemBytes 0x3012
00 00 00 12 00 0B 81 03 00 00 00 00 30 13 01 01 B1 04 00 00 00 02   S1F3  W=1  SystemBytes 0x3013

Byte 2 reads 81, 81, 82, 81. The top bit is the W-bit and the remaining seven are the Stream, so 82 is W-bit 1 plus Stream 2. Four requests means four transactions open at the same moment.

What came back:

RX 1  00 00 00 34 00 0B 01 04 00 00 00 00 30 10 ...   S1F4  SystemBytes 0x3010  (52 bytes)
RX 2  00 00 00 32 00 0B 01 02 00 00 00 00 30 11 ...   S1F2  SystemBytes 0x3011  (50 bytes)
RX 3  00 00 00 8B 00 0B 02 1E 00 00 00 00 30 12 ...   S2F30 SystemBytes 0x3012  (139 bytes)
RX 4  00 00 00 34 00 0B 01 04 00 00 00 00 30 13 ...   S1F4  SystemBytes 0x3013  (52 bytes)

Byte 2 on the replies is 01, 01, 02, 01 — the W-bit is down on all four. All four SystemBytes came back exactly as sent. Nothing is wrong here, which is why writing the host code against this shape alone walks you into the trap.

RX 1 and RX 4 are worth putting side by side. Same S1F3, same SVID, 2 ms apart, and both F8 values are the identical 40 50 5F 5C 28 F5 C2 8F. Asking the same thing 200 ms apart gave 65.0 and then 64.79. A polled value is a sample at an instant and the message carries no sample time, so even two values inside one batch give you no grounds to call them simultaneous.

One W=0 request shifts the whole batch

Now the same thing with the W-bit dropped on the second request only. The other three are unchanged.

TX (one write, 74 bytes)
00 00 00 12 00 0B 81 03 00 00 00 00 40 10 01 01 B1 04 00 00 00 02   S1F3  W=1  SystemBytes 0x4010
00 00 00 0A 00 0B 01 01 00 00 00 00 40 11                           S1F1  W=0  SystemBytes 0x4011
00 00 00 0C 00 0B 81 0B 00 00 00 00 40 12 01 00                     S1F11 W=1  SystemBytes 0x4012
00 00 00 12 00 0B 81 03 00 00 00 00 40 13 01 01 B1 04 00 00 00 02   S1F3  W=1  SystemBytes 0x4013

Byte 2 on the second line is 01. That is 81 with exactly one bit gone — the difference a human eye scanning a hex dump is least likely to catch.

RX 1  00 00 00 34 00 0B 01 04 00 00 00 00 40 10 ...   S1F4  SystemBytes 0x4010
RX 2  00 00 00 79 00 0B 01 0C 00 00 00 00 40 12 ...   S1F12 SystemBytes 0x4012
RX 3  00 00 00 34 00 0B 01 04 00 00 00 00 40 13 ...   S1F4  SystemBytes 0x4013
(waited 1.5 s more, nothing further — three replies for four requests)

0x4011 gets no reply. The equipment did nothing wrong. In SECS-II (SEMI E5) a W-bit of 0 means "do not send a reply", and the listener did precisely that. No error, no Stream 9 message.

This is where hosts split in two. Code that attaches each arriving reply to the head of its own request queue reads RX 2 (0x4012) as the answer to 0x4011. An S1F12 body lands in the S1F2 slot, so the parse either throws or, worse luck, quietly picks up the wrong value. And 0x4013 is left with no partner, so T3 fires 45 seconds later. The answer is on the wire and the log says timeout.

Code that pairs on SystemBytes notices nothing, because 0x4011 was never put in the transaction table in the first place — a request sent with W=0 is not something you are waiting on.

An unimplemented stream fills the slot with S7F0

Third capture. I slid an S7F19 — a message this listener does not implement — into the middle.

TX (one write, 58 bytes)
00 00 00 12 00 0B 81 03 00 00 00 00 50 10 01 01 B1 04 00 00 00 02   S1F3  W=1  SystemBytes 0x5010
00 00 00 0A 00 0B 87 13 00 00 00 00 50 11                           S7F19 W=1  SystemBytes 0x5011
00 00 00 12 00 0B 81 03 00 00 00 00 50 12 01 01 B1 04 00 00 00 02   S1F3  W=1  SystemBytes 0x5012

RX 1  00 00 00 34 00 0B 01 04 00 00 00 00 50 10 ...   S1F4  SystemBytes 0x5010  (52 bytes)
RX 2  00 00 00 0A 00 0B 07 00 00 00 00 00 50 11       S7F0  SystemBytes 0x5011  (10 bytes)
RX 3  00 00 00 34 00 0B 01 04 00 00 00 00 50 12 ...   S1F4  SystemBytes 0x5012  (52 bytes)

Take RX 2 apart: length 00 00 00 0A is 10, so it is header only with no body. Byte 2 07 is W-bit 0 plus Stream 7, and byte 3 00 is Function 0. An SxF0 aborts that transaction, and the part that matters is that it carries the aborted request's SystemBytes — 0x5011 is right there, so the host knows which request was refused.

Three replies for three requests, so the count matches and order-pairing survives this batch. Surviving is the worse outcome. It leaves a bug that only fires on batches containing a W=0 request, and the reproduction condition becomes "the day someone put a fire-and-forget S1F1 in the polling set".

This listener's ordering is not a guarantee

All three captures came back in request order, and the listener's code says why. It walks the receive buffer in a while loop, cuts out one complete frame and writes the reply on the spot. Serial handling inside a single socket event has no way to reorder.

That is this tool's implementation, not a property HSMS guarantees. SEMI E37 neither limits how many transactions may be open at once nor requires replies to come back in request order. On real equipment, where per-message application work differs — an S7 that has to read a recipe off disk open alongside an S1F3 that reads a value out of memory — the later request coming back first is the natural outcome. I have not verified a capture of out-of-order replies from real equipment. What I verified here stops at this: the header carries a field that makes pairing work without depending on order at all.

T3 is the same story. T3 is the reply timeout for one data message sent with the W-bit set, default 45 seconds with a 1–120 range. Four open transactions means four timers. Build it as one global timer started at "request sent" and the first transaction to finish turns off somebody else's timer.

What to put in the host code

  • The transaction table key is SessionID plus SystemBytes. Not arrival order, not the head of a queue.
  • A request sent with W=0 does not go in the table. Put it there and it will time out on T3, every time.
  • On an SxF0, close that SystemBytes as failed. Pass it to the next reply instead and everything after it shifts.
  • One timer per open transaction.
  • Allocate SystemBytes so they do not collide on that connection. Holding four open at once raises the odds, and once two collide there is nothing left to tell the replies apart. The allocation rules are in why SystemBytes uniqueness is your host's job.
  • Do not count replies. In two of these three captures the request count and the reply count differed.

Gluing four requests into one write() is not itself the problem. The kernel sends them as one segment and the receiver counts four frames — framing on the length prefix is pulled apart separately in one TCP read is not one HSMS message. What breaks is the pairing, not the transport.

To reproduce these bytes, start the passive listener on the SECS/GEM simulator and open a socket at it. Dropping one W-bit takes a minute, and that minute is one fewer "the equipment is not answering" ticket.