HSMS Linktest Timeout While Data Still Flows: Close the Socket Anyway
A swallowed Linktest.req burns T6 while an S1F1 on the same socket answers with a full S1F2 in 0.3 ms, and the host should close the connection anyway.
Articles
Field-oriented references for HMI/SCADA systems, industrial communication, alarms, tags, historians, and project delivery.
A swallowed Linktest.req burns T6 while an S1F1 on the same socket answers with a full S1F2 in 0.3 ms, and the host should close the connection anyway.
An S1F4 body carries no SVIDs at all. Real S1F11/S1F12 bytes showing where a host learns SVID numbers and names, and how S2F29 differs.
SELECTED is E37, COMMUNICATING is E30. Real bytes showing an S1F3 sent before S1F13/S1F14 aborted as S1F0, and the same S1F3 answered after.
You asked S2F13 for one ECID and S2F14 came back with two values. Real S2F29 and S2F13 bytes, and the length check host code must run.
Your report is defined but S6F11 never arrives. Real S2F33, S2F35 and S2F37 captures from two equipments, and what a non-zero LRACK tells the host.
Equipment clock drift reorders S6F11 events while HSMS stays green. Reading the clock with S2F17, setting it with S2F31, and why TIACK 0 proves nothing.
S2F41 START rejected with HCACK 2 while the equipment screen reads ONLINE. Separating the SEMI E30 communication state from the three control states, with a real HSMS capture showing the byte where HCACK flips from 2 to 0.
Run all ten SEMI E30 startup steps — S1F13 through S2F31 — with one npx command against a simulator, with the equipment-side wire capture of the run, what each PASS proves, and what a refused Select actually prints.
Five bad SECS-II messages pushed at a live HSMS listener. One drew a real S9F7, the rest came back as SxF0 aborts. What SEMI E5 stream 9 means and how to read MHEAD.
A real capture where two HSMS requests share one SystemBytes and the replies come back byte-identical, plus the allocation rules that prevent it.
One log line says timeout. A real HSMS capture separates T3 from T6, pins the SEMI E37 defaults (T3 45 s, T6 5 s, T7 10 s, T8 5 s) and shows the two timers your host has to enforce itself.
A 300-byte PPBODY is 330 bytes on the socket. A real S7F1 and S7F3 capture, and what a host loses by skipping the grant step.
Connection refused in 3 ms, or a TCP session where nothing ever arrives. Real captures that tell the two HSMS mode mistakes apart.
The HSMS session is SELECTED and S1F13 draws only a T3 timeout, while Linktest answers on the same socket in 0 ms. A capture, and two state machines.
An S1F3 body byte by byte: the item header packs format code and length-byte count, and one wrong length silences the whole connection.
A capture where Deselect.req, Reject.req and two unassigned STypes all draw silence, while Linktest answers instantly and the session stays SELECTED.
A capture of two host connections into one passive HSMS listener. Both get Select Status 0, and the first is never told the second arrived.
S1F1 sent, no S1F2, T3 expires — and the tool did nothing wrong. Captures of Header Byte 2 showing what packing the W-bit and Stream together costs.
A host that only drops replies under load is usually a parser treating one recv() as one message. Four captures: two messages in one write, one split, a body short of its length field, and a prefix claiming 1 MiB.
The host shut down but the tool still shows SELECTED. Two captures side by side: a session ended with Separate.req, and one where only the socket closed.
In HSMS, alive and connected are different states. A real capture of Linktest and S1F13 sent before Select ever goes out, the Reject.req reason 4 that SEMI E37 asks for there, and why the timer that fires first is T6, not T7.
Real Select.rsp captures — accepted, refused with SEMI E37 Select Status 1/2/3, and answered under a different SessionID — and where 'select failed' loses the detail.
Real captures of the three ways a Select gets no usable answer — silence, mismatched SystemBytes, and a wrong SessionID — and the SEMI E37 timers T5, T6 and T7 that decide what your log shows.
A real HSMS capture from 127.0.0.1:5501: request and mismatched Select.rsp decoded byte for byte, what a host's SystemBytes-keyed pending reply table does with the frame, and why the timer that fires is T6 and not T3.
Bringing WirelessHART gateway data into SCADA safely: update rate versus polling, stale-value detection, Modbus mapping pitfalls and join troubleshooting.
Remote notification is a delivery channel, not an alarm system. Picking which alarms page out, escalation timers that match humans, and proving delivery.
Pressure and temperature compensation for gas flow: which meters need it, the absolute-pressure trap, where the math belongs, and a failed P transmitter.
Building a strapping table into SCADA: segment geometry, interpolation spacing, where the calculation runs, and the density trap in hydrostatic level.
A level loop still cycling after two retunes is usually the valve. Reading stiction off PV/OP trends, proving it with a step test, and monitoring it.
Turning two limit switches and a travel timer into honest valve state: transit handling, discrepancy alarms, and catching a slowing actuator early.
Sizing polls, timeouts and keepalives for SCADA over LTE: the byte arithmetic, carrier NAT timeouts, report-by-exception, and why nobody dials in.
In ratio control one stream sets the pace and you trim the other. Where the multiplier goes, why actual ratio drifts, and starting up without off-spec.
One PID driving a heating and a cooling valve stalls or hunts right at the changeover. Setting the split point, killing overlap, and the gain jump.
Cascade loops fail when the two are tuned in the wrong order or the slave's limits never reach the master. Commissioning them so they stop fighting.
How IEC 62439-3 PRP and HSR deliver bumpless redundancy for SCADA and IEC 61850, why a failed LAN goes unnoticed, and what to monitor at commissioning.
Orifice DP goes as the square of flow, so someone takes the root — transmitter, PLC or SCADA. Do it twice, or skip the low-flow cutoff, and it lies.
NTP holds most SCADA clocks within milliseconds, which is plenty — until you need 1 ms sequence-of-events across substations. Where PTP earns its cost.
Where linearization and cold-junction compensation happen between sensor and SCADA tag, and why the wrong place quietly biases the reading you trust.
Smoothing a jittery 4-20 mA reading with first-order and moving-average filters, where the filter belongs, and why the wrong tag delays your alarm.
Designing rate-of-rise and rate-of-fall alarms: window length, noise filtering, units, and the quality gating that stops them screaming on comms recovery.
Changing ladder logic in a running PLC safely: online edit versus download, scan-order traps, seal-in glitches, forces, and a pre-edit checklist.
How a driver discovers BACnet devices, reads present-value and writes commandable objects: object identifiers, the 16-slot priority array, COV and BBMD.
How GOOSE stNum/sqNum counters, allowedToLive and ConfRev decide whether SCADA sees a breaker trip in milliseconds or never — and proving it is alive.
How BRCB and URCB blocks feed substation data to SCADA: datasets, trigger options, buffer time, and the reservation fights that starve a client.
DNP3 double-bit inputs (groups 3 and 4) encode 52a/52b contacts as four states, killing the OPEN/CLOSED flicker single-bit status shows mid-travel.
Commissioning DNP3 Secure Authentication (SAv5, IEEE 1815-2012): which functions to mark critical, aggressive mode on slow links, and key mismatches.
Reset windup is why a well-tuned loop overshoots after saturation or a manual-to-auto swap. What anti-windup and external reset actually fix in a PLC.
Why OPC Classic (DA/HDA) links fail across DCOM — RPC endpoint mapper on 135, dynamic port ranges, server identity — and the fastest way through it.
A PLC exposes a UDT over OPC UA and your client shows a ByteString. What an ExtensionObject actually holds, and how clients decode structured DataTypes.
Pulling ControlLogix and CompactLogix tags over EtherNet/IP: connected versus unconnected CIP, RPI, connection budgets, and why array reads win.
TCP to port 2404 is up, the RTU is reachable, every point still stale. It is almost always STARTDT and the general interrogation nobody sent.
The two IIN octets in every DNP3 response report restart, lost events, missing time and local control. Reading them, and clearing the ones that stick.
How to pull HART secondary variables and diagnostics into SCADA beyond the 4-20 mA loop, using multiplexers, HART-enabled I/O, and WirelessHART gateways.
Building lead/lag/standby pump rotation so duty sharing, standby failover and role swaps stay predictable — and visible to the operator at 3am.
A VFD is commanded through a packed control word and reports through a status word. Map either wrong and it won't start, or it clears a latched fault.
The deadband decides whether a change becomes an event; the Group 32 variation decides its size. Set both wrong and you flood the buffer or lose time.
A hot-standby pair earns its cost only if switchover is proven and sync loss is alarmed. What CPU redundancy covers, and testing both directions.
Commissioning a serial-to-Ethernet device server carrying Modbus RTU: operating mode, packing timers, TCP framing, idle timeouts and half-duplex control.
Running loop checks from field instrument through PLC to SCADA so every 4-20 mA point is proven end to end, and swapped transmitters surface early.
Computing usage across a wrapping counter without inventing phantom flow: register width, correcting the delta once, and rollover versus meter reset.
How drivers group Modbus registers into block reads, why gap tolerance and block size matter, and how one unmapped address takes down a whole block.
Sequence numbers and the Republish service turn a dropped notification into zero data loss — but only if the client watches them and the queue is sized.
A broken wire or saturated loop still scales to a plausible value. Detecting over-range, under-range and stuck 4-20 mA inputs before an operator acts.
Class 0/1/2/3 assignment, deadbands, unsolicited responses, event buffers, and the IIN bits that decide whether a master survives a comms gap.
A DNP3 link that stalls or logs CRC errors gets fixed one layer down: addressing, confirmed service, retry and timeout tuning on radio and serial.
Select-before-operate versus direct operate, CROB trip/close codes, the select timeout that bites on slow links, and verifying feedback not the response.
Why outstations stamp their own events, how the Need Time (IIN1.4) bit and delay-corrected clock writes work, and how bad sync turns an SOE log to fiction.
Designing and testing HMI command bits, acknowledgements and startup recovery so a stale command can't fire after a PLC or controller restart.
Running report-by-exception and change-of-state — DNP3 events, deadbands, heartbeats, integrity polls — without losing short events or stale values.
Alarm help text is written once during rationalization and never read again. What belongs in cause, consequence and response for a 02:00 operator.
During a flood a newest-first summary pushes the critical alarm off screen, and a filter left on erases it. Picking a default sort that holds.
Practical HMI notes for using color, shape, text, and state rules so operator screens stay readable during abnormal conditions.
Held values, dropped quality codes and late samples inserted at arrival time each make a trend lie after an outage. Working out which one you have.
Intermittent bad quality usually starts at a switch port. Which SNMP objects are worth polling, why counters only mean anything as rates, and trap gaps.
A Modbus TCP-to-RTU gateway can answer ping and still drop requests. Finding the queue saturation behind stale tags, and tuning the driver for it.
Browse returns success but the tag import is short. The OPC UA Part 4 v1.05 continuation point rules, MaxBrowseContinuationPoints, and how to find the cause.
Commissioning OPC UA PubSub (Part 14, UADP over UDP) without shipping stale values as live: dataset contracts, VLAN checks, and reboot-only failures.
Making an HMI screen state its own identity: ISA-101.01 display hierarchy as the breadcrumb skeleton, command labels that show scope, and testing copied screens.
A sequence screen showing only a step number sends the operator hunting across four displays. Separate state, step, and hold reason, and latch the reason.
How to choose Modbus RTU baud rate, silent interval (t3.5), and turnaround delay so slow serial devices get a fair chance to reply on a SCADA link.
MQTT arrival order is not process order. Why the DUP flag never catches the duplicates that hurt, Sparkplug's 8-bit seq wrap, and the MQTT 5.0 properties that fix stale retained state.
Shared subscriptions split an MQTT stream across workers — and will quietly halve your historian feed and scramble trend order if you point them wrong.
Estimating the real load behind site/#, and the MQTT 5.0 features that actually contain it: Retain Handling, shared subscriptions, Receive Maximum, SUBACK reason codes.
AccessLevel vs UserAccessLevel vs WriteMask: why a setpoint that browses writable returns BadUserAccessDenied, and testing with the production identity.
A maintenance bypass is an operating state, not a comment field. Tag naming, ISA-18.2 out-of-service handling, expiry rules, and the audit trail.
How historians hide communication gaps behind interpolated lines and held values, and the quality rules that keep bad data out of a shift report.
Wiring horn silence to the acknowledge tag corrupts the event log from day one. Splitting silence, acknowledge, reset and shelve under ISA-18.2.
How to split HMI button logic into authorization, mode, permissives, interlocks and quality — and show operators the actual reason a command is blocked.
Popup context passing for reusable faceplates: one stable equipment key, command binding that fails closed, and telling a broken binding from bad quality.
NodeIds regenerate, browse paths move, DisplayName is a label. Picking the reference your HMI and historian store so firmware doesn't blank a screen.
Choosing DataChangeFilter settings — absolute versus percent deadband, DataChangeTrigger, queue size — so a filter cuts noise, not real movement.
Repairing a historian gap with OPC UA HistoryRead: scoping the window, letting returnBounds handle edges, keeping Bad quality, and paging safely.
Take trend axis defaults from OPC UA EURange, not 0-100. Clamping at 100% hides NAMUR NE 43 fault currents that land at -2.5% and 106.25%, plus alarm lines and commissioning checks.
Recording calibration, manual mode and bypass windows as event records separate from the sampled stream, joined back by asset ID and time window.
A commissioning sequence for RSTP, MRP and PRP rings that measures what operators see: convergence time, dropped subscriptions, half-open sockets.
Modbus TCP devices cap simultaneous connections lower than you think. Inventorying clients, reading real socket behaviour, and protecting the HMI.
How a mishandled transaction ID lets a Modbus TCP gateway hand your SCADA driver the wrong register's value — and how to prove it with a capture.
Duplicate client IDs, wildcard rules nobody owns, retained commands that re-fire, and denied publishes that make no sound. What the MQTT spec actually guarantees at the edge.
The secure channel, the session and the subscription each expire on their own clock. Which one fired tells you to blame the firewall, load or keepalive.
Using StatusCode severity bits, SourceTimestamp and ServerTimestamp to catch frozen data, cached gateway values and clock drift a value-only HMI hides.
Why a flat 1-second scan overloads PLCs and gateways, and how scan classes and load shedding keep the data operators act on fast under stress.
Building an alarm priority matrix that ranks by consequence and operator response time, with the EEMUA 191 targets that keep high priority rare.
Exception and periodic collection store different truths. Picking per tag so a 300 ms interlock survives, and proving the choice during commissioning.
Rework quietly breaks linear MES route models. How to bind visit numbers, dispositions, and equipment events to the right route step instance.
Modbus has no standard endianness for 32-bit values, so two registers can decode to garbage. Proving byte and word order on site instead of guessing.
Versioning MQTT telemetry payloads so historians, HMI clients, MES connectors and analytics survive a field change instead of breaking on one.
Client won't connect though ping and port 4840 are open? Usually an endpoint URL, hostname SAN or ApplicationUri mismatch. Reading the status codes.
Standing alarms are acknowledged but still active. Reviewing them against the ISA-18.2 stale-alarm metric before they become summary wallpaper.
Historians key archive data to a point ID or a name string, so a rename can split a tag's history in two. What to verify before operators find the gap.
Where to tap, what to filter, and how to read the first thirty seconds of a capture when comms drop and recover before anyone can get to a keyboard.
The host wrote the EC, the tool returned EAC = 0, and nine lots ran on the old value. SEMI E5 EAC codes, a real S2F13 readback on the wire, the S2F29 namelist, and the SAT matrix.
A failed transmitter pegs a tag at full scale for eleven hours. Correcting the record with an audit trail that survives review, raw sample intact.
A one-scan reject pulse lives for 20 ms. Poll every 250 ms and SCADA never sees it. Matching polling rates to PLC scan time, driver and historian.
SEMI E30 spooling only covers streams the host enabled with S2F43. SPOOL LOAD vs UNLOAD, the S2F44 RSPACK and STRACK codes, S6F24 RSDA, and the S2F43 bytes on a real socket.
HMI, historian, PLC log and alarm list stamp one event seconds apart. Where each timestamp is born, and which one to trust during an incident.
The shift report says 1,738 good parts and the HMI says 1,742. Where to look first: counter style, historian deltas, S6F11 event time, or the spool.
Commissioning a Modbus RTU link the right way: prove polarity, termination, biasing and grounding before you ever open the register map.
Commissioning OPC UA Alarm & Condition subscriptions: event filters, the Retain flag, ConditionRefresh, and the failures that appear on reconnect.
A green redundancy icon proves the heartbeat works and nothing else. What to check on clients, alarms, historian and background jobs during a failover.
Averaging old process data hides trips and flattens peaks. Classifying tags, picking a method per class, and testing a purge job before production.
Why setpoint entry fails on the acknowledgment path, splitting display limits from PLC hard limits, and the tests that catch silent rounding.
A staged plan for rotating MQTT broker, client and CA certificates without breaking publishers, store-and-forward queues or Sparkplug birth sequences.
How OPC UA separates certificate trust, user identity and Part 18 role mapping — and commissioning writes so the wrong session can't move a setpoint.
A forced valve and a substituted flow reading are different overrides, and one Manual bit can't tell them apart. Logging that survives a server restart.
The Modbus unit identifier is 'useless' on TCP until a gateway is involved. How routing works behind serial gateways, and finding the slave that answered.
How reverse connect flips the TCP direction through a cell firewall, and the certificate, endpoint and diagnostic traps that still bite afterward.
OPC DA quality words, OPC UA StatusCodes, DNP3 flags — and how to display, alarm and historize them so a dead signal never reads as a valid value.
Split selected, downloaded and active recipe state into separate tags, compare what the PLC echoes against what you sent, and give rejections a reason.
Running a restore drill: what falls out of the backup set, the OPC UA certificate traps, RPO and retention per artefact, and a measured RTO against the target.
How monitored item queues, discard policy and the Overflow status bit decide whether your client sees every fast tag change or only the latest one.
Nobody memorizes a service account password, so a 90-day policy buys little. What IEC 62443-3-3 SR 1.2 and NIST 800-63B actually ask for.
One shared bit for acknowledge and reset is how alarms vanish while the fault is live. Separating active, latch, ack, return-to-normal and reset.
Split a five-second faceplate into first-load and update delay, then work down through tag counts, scan classes, trend queries and PLC load.
Why an MQTT edge gateway loses data during a WAN outage, and how to buffer with source timestamps, sequence numbers and a replay policy that holds.
Reading SEMI E5 HCACK and CPACK codes against a real S2F41 wire capture: why HCACK=4 exists, E30 control-state gating, and a retry that won't fire a second START.
Event frames pay off only if their boundaries hold. Where to take start and end triggers, what to capture as attributes, and frames that never close.
Where lot traceability breaks when built from equipment events: missing lot context, GEM event time, spool replay order, and split/merge genealogy.
A Modbus TCP write response says the transaction was accepted, not that the motor moved. Command discipline: permissives, readback, and the retry trap.
An OPC UA method call can return Good while the machine never moves. Wiring HMI calls so operators see acceptance, execution and a real failure reason.
A calculated tag quietly becomes an official number. Checking formula, quality rules, period boundaries and counter resets before reports depend on it.
Filtering an alarm banner is not suppressing an alarm, and ISA-18.2 draws that line for a reason. What the default view owes the operator during a flood.
Designing SCADA remote access against IEC 62443-3-3 SR 5.2 and SR 1.13, testing the denied paths at commissioning, and the port and timeout numbers that break it.
One screen says the pump is stopped, the report says faulted. The fix is a derived status tag with a written priority order that every client shares.
Setting on-delays, off-delays, debounce and deadband from process behaviour and ISA-18.2 — filtering nuisance alarms without hiding the first warning.
Alarm floods bury the initiating cause. Preserving first-out, trusting sequence-of-events timestamps, and stopping comms failures becoming floods.
How to drive HMI state from feedback instead of command bits, set timeouts from real actuator travel, and show operators why a command was rejected.
Linking SECS/GEM collection events with S2F33/S2F35/S2F37, why report content empties after a tool restart, and proving S6F11 matches the real sequence.
Commissioning SCADA firewall rules from data flows rather than port lists — the supporting services that break at cutover, and proving each rule for real.
A field-tested procedure for OPC UA redundant failover: ServiceLevel, subscription recovery, certificates, and catching stale data before operators do.
Why a client browses fine with security off but drops the secure channel, and how to fix trust stores, endpoints and certificate names in a project.
A field walkthrough for analog scaling: ranges, engineering units, Modbus word order, NAMUR thresholds — so PLC, SCADA, HMI and historian agree.
How historian timestamps go wrong — clock drift, NTP versus PTP, OPC UA source time, late data, DST — and catching it before a report dispute does.
Setting Modbus TCP poll rates, timeouts, retries and scan groups so one stalled device or saturated gateway can't drag the whole driver down with it.
Practical notes for choosing MQTT QoS levels, retained messages, clean sessions, and duplicate handling in SCADA and industrial telemetry projects.
How sampling interval, publishing interval, queue size, deadband and the KeepAlive/Lifetime pair decide whether your client and historian see the process.
Store-and-forward hands you hours-old samples when a link recovers. Which timestamp wins, what a late value may overwrite, and how to mark it.
How to show permissives, interlocks, inhibits, bypasses, command readiness, and first-out causes on practical HMI screens.
Reading Modbus exception responses on the wire, telling a refused request from a dead link, and chasing illegal-address and gateway causes.
After a PLC download the session opens and every item returns Bad_NodeIdUnknown. How namespace indexes and client caches shift, and what to check first.
Running alarm rationalization sessions that produce usable priorities, causes and operator response guidance instead of another unread spreadsheet.
Designing downtime reason codes operators pick correctly under pressure: prompt timing, auto-coding from PackML state, and the fields reports need.
Using Last Will, birth messages and retained state so a SCADA screen shows offline, stale and healthy as three things — not one green icon that lies.
A practical checklist for renewing OPC UA application certificates without breaking SCADA clients, historians, gateways, or production HMI connections.
Picking historian deadbands and compression from the instrument and the process, not HMI decimal places, so trends keep movement without storing noise.
Mode is who may command the equipment; state is what it is doing. SEMI E30's control state model splits them — here is how to split the HMI to match.
Modbus connects, unit IDs respond, polls complete — and the operator still reads a swapped float. Proving a register map before the historian trusts it.
A SCADA screen shows a believable number long after the device stopped updating. Heartbeat, watchdog and stale-data tags that make freshness visible.
Project notes about naming, alarms, screenshots, networking, commissioning, backups, and documentation habits that matter later.
Laying out Sparkplug B group, edge node and device IDs — plus birth certificates, aliases and STATE — so a rename doesn't break every subscription.
A layered checklist for troubleshooting SCADA communication problems from physical link to protocol behavior.
S5F1 carries ALCD, ALID and ALTX. Bit 8 of ALCD is set-versus-clear, and dropping it leaves an MES alarm list that never goes green. Plus the S5F5/S5F6 recovery after a host restart.
What an industrial historian does, what to plan before collecting data, and common mistakes in historian projects.
A practical checklist for commissioning OPC UA client connections from SCADA, HMI, historian, or gateway software to PLCs and automation servers.
Trend screens that answer the operator's real question during an upset: grouping by loop, choosing scales, and drawing a comms gap as a gap.
How raw SCADA, HMI and protocol values become useful machine state, process state, alarm state and control state, and where each one is decided.
A practical checklist for designing HMI faceplates that operators can use quickly, consistently, and safely during normal operation and troubleshooting.
A field checklist for commissioning managed industrial Ethernet switches used by SCADA servers, PLCs, remote I/O, drives, cameras, and protocol gateways.
How ISA-18.2 splits shelving, designed suppression, out-of-service and interlock bypass — and configuring them so every hidden alarm keeps an owner.
A commissioning checklist for validating HMI and SCADA screens, tags, alarms, trends, historians, users, backups and comms paths before handover.
A practical guide for creating readable, maintainable SCADA tag names across equipment, signals, alarms, trends, and reports.
A short hands-on walkthrough that turns one temperature signal into alarm state, switch action, HMI feedback, and an event log using the StateX tutorial.
Practical alarm design mistakes that create nuisance alarms, floods, operator confusion, and poor incident response.
How to think about OPC UA and MQTT in industrial systems without turning the comparison into a vendor argument.
A concise practical guide to Modbus TCP concepts: clients, servers, unit IDs, function codes, registers, polling, and mapping pitfalls.
A practical definition of SCADA tags and how they connect PLC addresses, alarms, trends, screens, and reports.
A field-oriented comparison of common industrial software layers and where their responsibilities overlap.
A practical explanation of SCADA systems: what they do, where they fit, and what to watch for on real projects.