← Articles
SCADA/8 min read/ views

Your Loop Isn't Badly Tuned — The Valve Is Sticking: Finding Stiction in Trend Data

A level loop still cycling after two retunes is usually the valve. Reading stiction off PV/OP trends, proving it with a step test, and monitoring it.

SCADAControlTroubleshootingTrendsHistorian

A tank level loop had been cycling for months with a period of about four minutes and a swing of roughly ±6% of span. Two different people had retuned it. The second retune halved the gain, which changed the period to about six minutes and made the swing slightly worse. By the time I saw the trend, the operators had given up and were running it in manual, nudging the valve by hand a few times a shift.

Pull up the trend and the shape tells you what's wrong before you touch a tuning constant. The level (PV) was a clean triangle wave — straight ramps up, straight ramps down, sharp corners. The controller output (OP) was almost a square wave, sitting flat then jumping. That pair of shapes is the signature of a valve that will not move until the controller pushes hard enough, and then moves too far.

What stiction actually is

Static friction in the valve packing and, less often, in the actuator. The stem is stationary and the friction holding it exceeds the force the actuator is applying. The positioner keeps integrating up the applied force until it wins, and then the stem breaks free and slips past the position it was aiming for — because kinetic friction is lower than static, so the moment it starts moving there's excess force.

Break it into two numbers you can measure:

  • Deadband — how much the input signal has to reverse direction before the stem moves at all. Friction plus any mechanical lost motion in the linkage.
  • Slip-jump (stickband) — how far the stem overshoots when it finally breaks loose.

The deadband is why the loop goes open-loop for a while. The slip-jump is why it overshoots when it closes again. Put an integrating controller around that and you get a limit cycle that no set of tuning constants will kill. That's the important part: with real stiction and any integral action at all, the loop will cycle. Tuning only changes the period and amplitude. If someone tells you they'll tune it out, they're going to spend a day proving they can't.

Distinguishing it from bad tuning and from a real disturbance

Three checks, in the order that costs the least:

Put the loop in manual and leave the output alone. If the oscillation stops within a couple of process time constants, the cycle is being generated inside the loop — valve, or tuning. If it keeps cycling with a frozen output, it's an external disturbance (an upstream loop, a batch cycle, a compressor loading and unloading) and the valve is innocent.

Look at the PV shape. Tuning-induced oscillation on a self-regulating process is roughly sinusoidal and it decays or grows smoothly. Stiction on an integrating process (level, pressure in a closed vessel) gives you triangles and sawtooths with corners. Non-linear shape means non-linear cause.

Change the gain and see what the amplitude does. Cut controller gain in half. Tuning-induced cycling settles down. Stiction cycling gets slower and typically no smaller, sometimes bigger — which is exactly what happened on my level loop and is why the second retune made it worse.

There's also the OP-versus-PV plot, which people call the phase plot or the mapping. Plot controller output on one axis and valve position (or PV, if that's all you have) on the other over a few cycles. A healthy valve traces something close to a line. A sticking valve traces a parallelogram — the flat sides are the stem not moving while the signal sweeps across the deadband. Choudhury's stiction detection work is built on exactly this, and once you've seen a couple you can spot it by eye in a trend client without any math.

The step test that settles the argument

Trends give you a strong suspicion. A step test gives you a number you can put in a work order.

Take the loop to manual and step the output down in decreasing increments — 5%, 2%, 1%, 0.5% — with a pause of a few time constants between steps, then reverse and walk it back up the same way. Trend the commanded output and the positioner's position feedback together. What you're looking for:

  • The smallest step in one direction that produces any stem movement at all. That's your resolution.
  • The signal change required after a direction reversal before the stem moves. That's your deadband.
  • Whether the stem lands where it was told or shoots past and settles back. That's your slip-jump.

ANSI/ISA-75.25.01 is the test procedure for control valve response measurement from step inputs, and ISA-TR75.25.02 covers the terminology and the dynamic performance parameters, if you want to specify this properly with a vendor rather than argue about "sticky". As a working rule, a modulating control valve on an important loop that can't resolve a 1% step, or that needs more than about 1% of signal reversal before it moves, is going to cause you trouble. Fast, tight loops want considerably better than that. Loose loops on a big tank will happily tolerate a few percent.

One field caveat: do this with the process running at normal pressure if you can. Packing friction is real, but so is the unbalanced force across the plug, and a valve that steps beautifully on a shutdown line can be much worse at 40 bar differential.

What SCADA should be recording

Most systems historize PV, SP, and OP for every loop and stop there. That's the single biggest reason stiction goes undiagnosed for years — without position feedback you can only infer the valve's behavior from the PV, which works but is slower and more arguable.

If the positioner has a 4–20 mA position transmitter or the valve is on a fieldbus that publishes actual travel, historize it. My default is a 0.25% deadband on the position tag, exception-based. It's a slow-moving analog on a well-behaved valve, so it costs almost nothing in storage, and the moment a loop starts misbehaving you have months of history showing whether the stem was following the signal.

With position feedback available, a cheap derived tag catches the problem before anyone complains:

  • Accumulate absolute change in OP over a rolling window (say five minutes).
  • Accumulate absolute change in position feedback over the same window.
  • When OP travel exceeds a few percent and position travel is near zero, flag it. That's the valve refusing to move while the controller works.

The reverse ratio is worth watching too. Total OP reversals per hour is a decent proxy for valve wear across a whole plant — sort your loops by it and the sticking valves and the badly tuned ones both float to the top, which is a useful list either way. Don't alarm on it. This is maintenance information, not an operator action, and putting it on the alarm summary just adds to the flood.

What to do about it before the valve comes off the line

You can't fix friction in software, but you can stop it from wrecking the process:

  • Reduce integral action rather than gain. Longer integral time means the controller pushes across the deadband more slowly, which lengthens the cycle period and usually shrinks amplitude. It makes the loop lazier on genuine upsets. That's the trade.
  • Check the positioner first, not the packing. A badly tuned or undersized positioner produces symptoms that look exactly like stiction, and it's an hour of work instead of a shutdown. Confirm the positioner's own gain and that it has enough air supply and volume for the actuator.
  • Look at packing history. Environmental low-emission packing sets are much tighter than the old graphite arrangements and are a common cause of a valve that got worse after its last overhaul. If the loop's performance stepped down right after a turnaround, that's your date stamp.
  • Stiction compensation exists in some DCS controllers — a small pulse added to the output to knock the stem loose. It works, and I'd still treat it as a bridge to the repair, not the repair. It puts extra cycles on the packing.

If the position feedback follows the output cleanly through the step test and the loop still cycles, stop looking at the valve. Next stop is the interaction between loops — two controllers fighting over the same flow, or a level loop cascaded to a flow loop whose own valve is the sticky one. And if the stem tracks perfectly but the process doesn't respond, you may have a plug that's parted company with the stem, which the positioner will happily report as a perfect valve right up until someone opens it up.