SCADA Data Stops Updating: What to Record Before Requesting Field Support

When SCADA data stops updating, the fastest path to resolution is usually documentation, not immediate intervention. Record what failed, when it failed, which tags and devices are affected, what the communication status and tag quality show, and what changed recently , before restarting anything. That information lets a support technician narrow the probable cause remotely and arrive with the right tools and spares.

Key Takeaways

  • A restart often clears the very evidence needed to identify the cause. Capture screenshots, logs, and status indicators first.
  • A frozen value with good quality points somewhere different than a value with bad or uncertain quality. The distinction changes where the investigation starts.
  • Live data working while trends show gaps is a historian or archiving issue, not necessarily a field communication failure.
  • Protocol matters for evidence: DNP3 devices buffer time-stamped events and may backfill after communication is restored; a polled Modbus link generally does not.
  • Most SCADA data failures trace to a recent change. Reconstruct the 72 hours before the failure before assuming hardware failed.

Data stopping on a SCADA screen is a symptom, not a diagnosis. The same symptom can come from a stopped PLC scan, a failed radio path, an expired OPC server license, a firewall rule change, a full historian disk, or a single mis-mapped register. The reason so many of these calls become multi-trip service events is that the first response, restarting the server, cycling the driver, or rebooting the RTU, removes the information needed to tell those causes apart.

This article covers what to record while the failure is still visible, how to classify the failure, and what to send with a support request so the person responding can work from facts instead of guesses. It is a reporting and evidence-capture guide, not a repair procedure. Any work involving energized equipment, panel entry, or field device isolation requires qualified personnel and your facility’s hazardous-energy procedures.

What to Record When SCADA Data Stops Updating

Why Restarting First Makes Diagnosis Harder

A restart of a SCADA server, communication driver, OPC server, or RTU resets counters, clears cached status, drops the current connection state, and in many systems rolls over diagnostic buffers. If the restart happens to restore data, the underlying cause remains unknown, and an intermittent fault that clears on restart will almost certainly return, usually without warning.

This matters most for intermittent failures. A link that drops once a week is far harder to characterize after six restarts than after one well-documented event. If the process is stable and operators retain the visibility they need to run safely, there is usually time to spend five minutes capturing evidence. If the loss of visibility creates a safety or process risk, follow your facility’s procedures for operating without SCADA indication first; evidence capture never takes priority over safe operation.

Screenshot SCADA Screens, Tag Values, and Timestamps

Capture the affected process graphic showing the stale values, and include anything on screen that shows time: the SCADA clock, a last-update or last-scan timestamp, a communication status block, a heartbeat or watchdog tag. A screenshot with no visible timestamp loses most of its diagnostic value.

Take a second screenshot of an area of the system that is still working normally. The contrast between a working section and a failed section is one of the quickest ways to bound the problem.

Export Alarm, Event, and Communication Logs

Export the alarm and event log for a window that starts well before the failure, several hours earlier at minimum, longer for an intermittent problem. Many systems overwrite or roll these logs on a fixed schedule, so exporting early protects the record.

Communication logs usually live separately from process alarms: the driver’s own log, the OPC server log, and in some architectures a dedicated gateway or front-end processor log. Export each one you have access to. A communication driver log that shows the exact timeout sequence is often more useful than the entire process alarm history.

Record Driver, OPC, PLC, and RTU Status Before Any Change

Write down or photograph the state of each element in the chain while it is still in the failed condition:

  • Communication driver or channel status (connected, disconnected, error, retry count)
  • OPC server connection state and the quality indication on affected items
  • PLC or RTU front-panel LEDs run, program, fault, and network link and activity indicators
  • Network switch port link and activity indication, where safely visible
  • Radio or cellular signal strength and connection status at remote sites

Read indicators that are visible from outside the enclosure. Do not open energized panels or probe conductors to collect this information unless you are qualified to do so and the work is covered by your facility’s electrical safety program.

Determine What Kind of SCADA Data Failure Is Occurring

Classifying the failure narrows the data path before anyone drives to site. Five questions do most of the work.

One Tag vs. Multiple Tags Not Updating

A single stale tag while everything else on the same device updates normally rarely indicates a communication failure; the link is demonstrably working. More likely candidates include a tag configured against the wrong register or address, a change in the PLC program that stopped writing that value, an instrument signal problem upstream of the PLC, or a value that genuinely is not changing.

When every tag from a device goes stale together, the problem sits at the device or the path to it. Record which it is, and list the specific tags either way.

One PLC or RTU vs. Multiple Devices

If one device is silent and others on the same network segment and the same driver are fine, attention moves to that device, its power, and its local connection. If several devices fail together, look for what they share: a switch, a radio repeater, a network segment, a driver instance, a firewall rule, or a common power source.

Note which devices are affected and which comparable devices are not. Both lists matter.

One Site vs. Multiple Sites

For distributed systems, determine whether the failure follows a geographic pattern or a network pattern. Multiple remote sites dropping points toward a shared communication path, a cellular carrier or APN issue, a VPN or gateway problem, or a change at the central end. A single remote site failing alone points to that site’s power, radio path, modem, or local equipment.

Stale Live Data With Normal Communication Status

This is one of the most informative conditions to report. If the driver shows a healthy connection, the device responds, and tag quality reads good, but values are frozen, the communication path is probably intact, and the data itself has stopped changing at the source.

Possible explanations include a PLC placed in program or stopped mode, logic that stopped executing the routine updating those registers, a device left in a forced or simulation state, or a field instrument output that has locked at its last value. Some drivers also continue serving the last successfully read value under certain configurations, which can make a dead link look healthy on the screen. Each of these requires verification; none can be confirmed from the symptom alone.

Live Data Working but Historical Data Missing

If an operator sees current values updating but trends show a gap, the failure is downstream of the SCADA runtime, in the historian, data logger, or archiving layer. Common contributors include a stopped historian service, a full or nearly full disk, a store-and-forward buffer that filled during an outage, a tag not configured for historization, a licensing or tag-count limit, or deadband and compression settings that suppress logging when values barely change.

Record whether the gap has a clean start and end, and whether data resumed on its own. A gap that backfilled later suggests buffering recovered; a permanent gap suggests data was never written.

Record When the Data Stopped Updating

Date and Time of the First Observed Failure

Record two separate times: when the data actually stopped, read from the last good timestamp in the system, and when someone first noticed. These are often hours apart, and the difference tells support how much log history to examine.

Use the plant’s standard time reference and note the time zone explicitly. Daylight saving transitions and mixed UTC/local configurations are a recurring source of confusion in distributed systems.

Continuous vs. Intermittent Data Loss

A failure that has been continuous since a fixed moment and one that comes and goes are investigated differently. For intermittent loss, record the pattern as precisely as you can: how often, how long each outage lasts, and whether it correlates with time of day, weather, equipment starting, a scheduled backup, or a shift change.

A pattern tied to a recurring event is a strong lead. Intermittent problems with no apparent pattern usually require diagnostic logging over time rather than a single site visit, and saying so up front sets realistic expectations for the response.

Comparing SCADA, PLC, and RTU Timestamps

Where the devices support it, compare the timestamp on the last good value at the SCADA level with the device’s own clock or last-event time. A mismatch can reveal clock drift, an NTP synchronization problem, or a buffering layer holding data.

Protocol behavior matters here. DNP3 outstations can time-stamp events locally and buffer them, so data may appear in the historian after communication returns, carrying the original field time. Polled Modbus registers carry no inherent timestamp; the SCADA applies its own at read time, so a Modbus-based system generally has no field-side record of what happened during the gap. Knowing which applies tells support whether recoverable data exists.

Record the Affected Tags and Equipment

Tag Names, Descriptions, and Register Mappings

Provide exact tag names as configured, not operator shorthand. Include the description, the device the tag belongs to, and the register, address, or item path if you can see it in the configuration without making changes. Exported tag lists are ideal when available.

PLC or RTU Names and Site Locations

List affected devices by their configured name, their model where known, and their physical location. For remote assets, include the site identifier used in your system and how the site communicates back: licensed radio, cellular, leased line, fiber, or satellite.

Affected Equipment or Process

State what the stale data represents operationally: a pump station, a compressor, a tank level, a flow meter, a motor starter. This tells support whether the issue touches a control function, an alarm, a regulatory or environmental measurement, or a view-only indication, which affects both urgency and the sequence of restoration.

Expected Values vs. Displayed Values

Record what the screen shows against what the process is actually doing, if you can verify it locally. A level frozen at 48% while the tank is visibly filling, a flow reading zero while the pump runs, or a value pinned at full scale all point in different directions. Frozen-at-last-value, zero, and full-scale behave differently and are worth distinguishing.

Tag Quality and Communication Status

Tag quality is one of the most diagnostic pieces of information in the whole report, and it is frequently left out. Record the quality indication on the affected tags along with any sub-status your system displays.

In OPC terms, good means the server believes the value is valid, bad means it is not usable, and uncertain means the value is being provided but its validity cannot be confirmed, a common state when a server is serving a last known value or operating in a degraded condition. Good quality with a frozen value and bad quality with a frozen value are different problems. Copy the exact wording or code your system displays rather than paraphrasing it.

Identify Where the Data Path Breaks

Data moves through a chain, and each link has its own failure modes. Recording where the break appears, even approximately, is more useful than reporting that “SCADA is down.”

Field Device to PLC or RTU

The first link is the instrument and its wiring to the I/O. Symptoms here are usually confined to individual points rather than whole devices: one analog input reading zero or full scale, a transmitter in a fault state, a discrete input stuck in one state.

Confirming this link typically requires loop checks, signal measurement, or instrument diagnostics by qualified personnel. If you suspect a field signal rather than a communication problem, record which point and what it reads, and route it as instrumentation work; our instrumentation and electrical services and SCADA support cover different parts of this chain, and saying which you suspect helps get the right technician assigned.

PLC or RTU to Industrial Network

Record the device’s run/fault indication and its network port link and activity status. A PLC in program or stopped mode, a failed communication module, a loose or damaged cable, or a serial converter problem all live here. For serial links, note the configured settings, baud rate, parity, data and stop bits, and the device address since a mismatch after any configuration change breaks communication immediately.

Network to Communication Driver or OPC Server

This link covers switches, routers, firewalls, VPN tunnels, and the radio or cellular path for remote sites. Record whether the device is reachable at the network level and whether the change coincided with any network work.

Keep diagnostics conservative on operational technology networks. Basic reachability checks are generally acceptable; broad port scanning or active discovery tools can disrupt control devices and should not be run on a live ICS network without planning and authorization. NIST SP 800-82 Rev. 3 covers why OT environments require different handling from standard IT networks.

Driver or OPC Server to SCADA

Record the driver or channel status, error and timeout counters, retry settings, and any messages in the driver log. Also note whether the OPC or driver service is running, whether it restarted recently, and whether any licensing warning is displayed. Expired or exceeded licensing is a recurring cause of data that stops cleanly at a precise moment with no network or device fault at all.

SCADA to Historian or Data Logger

This is the home for every trend and archiving issue. Record the historian service state, free disk space on the archive volume, any store-and-forward buffer status, and whether the affected tags are configured for historization. If trends are missing for some tags but not others on the same device, configuration is a more likely explanation than communication.

Record Diagnostic Messages and Error Conditions

SCADA Alarm and Event Messages

Copy the exact text of communication-related alarms and system events, with their timestamps. Verbatim wording matters; system messages are often specific enough to identify the failing component, and a paraphrase usually loses that specificity.

Communication Driver Errors and Timeouts

Record the error text, error code, and how often it repeats. Distinguish between a timeout (the request was sent, no response arrived), a connection refused or reset (something actively rejected it), and a protocol-level exception returned by the device (the device responded, but could not serve the request). These three point in notably different directions.

OPC Server Status and Quality Codes

Note the server’s connection state, the group or subscription status, and the specific status or quality codes on affected items. If the system shows OPC UA StatusCodes or OPC DA quality values, record them exactly as displayed.

PLC or RTU Fault and Run Status

Record front-panel indication: run, program, fault, battery, and network LEDs, plus any fault code shown on a display. Where your platform provides it without requiring a configuration change, note the processor mode and any logged major or minor fault.

Network Connectivity Test Results

If your site’s procedures permit basic connectivity testing, record what was tested, from where, and what the result was, including response times and packet loss, which matter on radio and cellular links where high latency and intermittent loss can cause timeouts even when the device is reachable. Note the source machine for each test; a device reachable from the IT network but not from the SCADA server is a routing or firewall finding, not a device failure.

Record Recent Changes That Could Have Caused the Failure

Most SCADA data failures follow a change. Reconstruct the period before the failure, typically the preceding 72 hours, longer if the system is only checked periodically, and ask everyone with access, including contractors and vendors.

PLC Program Downloads and Tag Changes

Note any program download, online edit, tag addition or rename, or address change. A renamed or re-addressed register breaks the SCADA mapping even though both the PLC and the network remain perfectly healthy. Record who performed it, when, and whether a backup of the previous program exists.

Firmware and Software Updates

Record firmware updates to PLCs, RTUs, radios, modems, or switches, and software updates to the SCADA platform, OPC server, drivers, or the underlying operating system. Operating system updates that restart services or change security settings are a frequent and easily overlooked cause.

SCADA or Driver Configuration Changes

Include changes to channel or device settings, polling rates, timeout and retry values, tag scaling, security settings, user accounts, and license files. A shortened timeout can break a marginal radio link that previously worked.

Network, Firewall, or Routing Changes

Record firewall rule changes, VLAN or subnet changes, IP address changes, switch replacement or reconfiguration, VPN certificate renewals, and carrier or SIM changes on cellular links. Network changes made for legitimate security reasons are a common cause of OT communication failures, particularly where OT traffic was not fully documented beforehand.

Power Outages, Equipment Trips, and Network Failures

Note outages, UPS events, generator transfers, breaker trips, lightning activity, and any equipment that tripped around the time of the failure. Equipment that lost power and restarted may have come back with a different IP address, in program mode, or with a cleared configuration, all of which present as a data failure afterward.

Document What You Already Tried

Troubleshooting Steps Already Performed

List every action taken, in order, with the time of each: services restarted, devices power-cycled, cables reseated, configurations changed, tests run. Include actions that produced no result. Unrecorded steps get repeated on site, and repeated steps consume time that was already paid for once.

What Changed or Did Not Change as a Result

Record the outcome of each action. “Restarted the OPC server at 14:10, data resumed for about 20 minutes, then went stale again” is a high-value observation; it indicates the fault recurs, bounds the interval, and rules out a permanent hardware failure in that component.

Why This Prevents Repeated Work on Site

A technician arriving without this record has to rebuild it before any new diagnostic work begins. With it, they can plan the sequence in advance, bring the right test equipment, and arrange the right site access and permits. For remote facilities, where a single trip can consume most of a day, that preparation has a direct effect on how quickly the issue is resolved.

SCADA Troubleshooting Considerations for Louisiana Industrial Sites

Remote Oil, Gas, and Process Facilities

Many Louisiana sites rely on SCADA for visibility into equipment that no one stands next to. When data stops updating at an unmanned location, the practical consequence is operating without current indication until someone travels there, which is why accurate remote characterization matters more here than at a facility where an operator can simply look at the equipment.

Site access is often the long pole: gate keys, safety orientation, escorts, hot work or permit requirements, and hazardous-area considerations all need arranging before a technician can work. Including access requirements in the initial support request avoids a wasted trip.

Cellular and Radio Communication at Remote Sites

Wireless links introduce failure modes that wired networks do not. Record signal strength readings over time rather than a single instantaneous value, and note anything that changed in the path, new construction, vegetation growth, an antenna that may have moved, or a repeater site with its own power or equipment issue.

Carrier-side changes also matter: SIM or plan changes, APN or private network configuration changes, IP addressing changes, and regional carrier maintenance can all interrupt telemetry with no fault at the site itself. If multiple remote sites failed at once and they share a carrier, record that.

Power, Weather, and Environmental Factors

Note recent storms, lightning, flooding, extended heat, or power events near the time of failure. Surge damage to communication equipment, radios, and I/O is a real possibility where exposed outdoor equipment is involved, and moisture intrusion in enclosures and connectors can produce intermittent faults that look like network problems.

Record the environmental context rather than concluding from it. Weather coinciding with a failure is useful evidence; it is not a diagnosis, and equipment should be inspected before any component is declared damaged.

What Local Field Support Needs to Arrive Prepared

The most useful support request contains the failure classification, the timeline, the affected tags and devices, the quality and status indications, the error messages, the change history, and what has already been tried, along with site access requirements and a point of contact who can authorize work.

Advanced Energy Services provides SCADA and automation support, control-system troubleshooting, and instrumentation and electrical services to industrial, energy, utility, and process facilities across Louisiana, Texas, and the Gulf Coast region from its base in Broussard, Louisiana. Sending the information in the checklist below with your request lets the response begin with analysis rather than rediscovery. You can reach the team through the SCADA and automation services page to discuss an active data failure or a recurring communication issue.

Frequently Asked Questions

Does restarting the SCADA server fix stale data?

Sometimes, but it is rarely the right first step. A restart can clear a hung driver or service and restore data temporarily, while also erasing the diagnostic state needed to identify why it hung. If the cause sits in the PLC, the field device, or the network, a server restart will not help at all. Capture evidence first, and treat a restart that “fixes” the problem as an unresolved intermittent fault rather than a resolution.

What does bad or uncertain tag quality mean?

Quality describes the server’s confidence in the value, independent of the value itself. Good indicates the value is believed valid. Bad indicates it should not be relied on , typically a lost connection, a device not responding, or a configuration error. Uncertain indicates a value is being supplied but cannot be confirmed as current or accurate, often because a last known value is being served or the source is partially degraded. Because a stale number can still display with good quality, quality and value have to be reported together.

Why does live data update but historical data stay missing?

Live values and historical records usually travel through different parts of the system. Current values come from the communication driver to the SCADA runtime; historical data is written separately by a historian or logging service. If the runtime is receiving data but the historian is stopped, out of disk space, out of licensed tag capacity, or not configured to log those specific tags, trends will show gaps while screens update normally. This points the investigation at the archiving layer, not the field.

How do I tell a network problem from a PLC problem?

Compare behavior across devices. If several devices on the same network path fail together while devices on other paths are fine, the shared network element is the stronger candidate. If one device is unreachable while its neighbors on the same switch and subnet respond normally, attention shifts to that device, its power, and its local connection. A device that responds at the network level but returns errors or frozen values suggests the link is intact and the problem lies in the device’s state or program. None of these is conclusive on its own; they narrow where testing should begin.

What information should I send with a field support request?

At minimum: when the data stopped and when it was noticed; which tags, devices, and sites are affected and which comparable ones are not; the quality and communication status on affected tags; exact error messages from the SCADA, driver, and OPC logs; any changes made in the preceding days; what you have already tried and what resulted; and site access requirements. Screenshots and exported logs should accompany the written summary. The checklist below covers the full set.

SCADA Data Failure Reporting Checklist

Work through this before contacting support, and send it with your request.

Evidence captured before any changes

  • Screenshots of affected screens showing stale values and visible timestamps
  • Screenshot of a comparable section still working normally
  • Exported alarm, event, and communication logs covering before and after the failure
  • Photographs or notes of PLC, RTU, switch, and radio indicator status

Failure classification

  • Single tag, multiple tags, single device, multiple devices, single site, or multiple sites
  • Values stale with good quality, or bad/uncertain quality
  • Live data only, historical data only, or both

Timeline

  • Last good data timestamp and time of discovery, with time zone
  • Continuous or intermittent, with the observed pattern if intermittent
  • Any discrepancy between SCADA and device timestamps

Affected tags and equipment

  • Exact tag names, descriptions, and register or address mappings
  • Device names, models, and site locations
  • Equipment or process affected, and whether control, alarms, or regulatory measurements are involved
  • Displayed values against expected values
  • Quality and communication status, copied exactly as displayed

Diagnostic messages

  • Verbatim SCADA alarm and event text with timestamps
  • Driver error codes and timeout messages, with repetition frequency
  • OPC server status and item quality codes
  • PLC or RTU fault codes and processor mode
  • Connectivity test results, including the source machine for each test

Recent changes (preceding 72 hours or more)

  • PLC program downloads, online edits, tag or address changes
  • Firmware, software, and operating system updates
  • SCADA, driver, or OPC configuration changes
  • Network, firewall, routing, VPN, or carrier changes
  • Power outages, trips, storms, and network failures

Actions already taken

  • Every step performed, in order, with times
  • The result of each step, including steps that changed nothing

Site and access

  • Location, access requirements, permits, escorts, and safety orientation
  • Site contact and the person authorized to approve work
  • Communication method at the site, radio, cellular, leased line, fiber, or satellite

Any step requiring panel entry, energized testing, or field device isolation should be performed only by qualified personnel under your facility’s electrical safety and hazardous-energy procedures. If loss of SCADA visibility affects safe operation of the process, address operating requirements first and treat the data failure as a separate workstream.

Scroll to Top