Host Engineering Forum
General Category => Do-more CPUs and Do-more Designer Software => Topic started by: JeffS on September 29, 2023, 05:50:26 PM
-
I have a few PLCs that are pretty consistently triggering the hardware watchdog timeout. Attached are the system logs for the two PLCs. Is there anything I can do to narrow down what could be causing this? For the Test BRX Maestro it has nothing wired up to any of the cards attached to the BRX PLC. On my test BRX I plan to unplug the connectors on every card so they don't have external power, just to see if that affects the hardware timeout occurrence.
Carnegie PLC: 2.9.6 OS, 1.1.0 Booter
Test PLC: 2.9.7 OS, 1.1.0 Booter
What can I do to help narrow this down to a particular card or whatever is actually causing the timeout?
Thanks,
Jeff
-
Normally, a hardware watchdog trigger is caused by something physical like:
- Radiated electrical noise (e.g. VFDs nearby?)
- I/O wiring/routing in parallel with high voltage
- No snubbers on relay I/O
- Dirty power and/or brown outs
- Changes/events to PLC program and/or the machine process
The System Data Word, DST385 ($WatchdogReboots) shows a count of how many times it has rebooted.
DST400-409 contain codes for us (Host Engineering) that may help narrow down what is happening. Can you post those?
-
Attached are some screen shots of the contents of DST385 and DST400-409. I have this happening on another PLC in the field as well at a separate location.
Where DST385 is 0 is because the PLC had dropped out of run after reaching 10, which each of these PLCs had all done at the time I logged in to get the data.
-
Attached are some screen shots of the contents of DST385 and DST400-409. I have this happening on another PLC in the field as well at a separate location.
Where DST385 is 0 is because the PLC had dropped out of run after reaching 10, which each of these PLCs had all done at the time I logged in to get the data.
The codes in DST400-409 look normal. But DST385 ($WatchdogReboots) = 0 is a bit troublesome.
Does this ever increment?
Also, since these all look normal (except for the DST385), what are these values?
- DST410 ($ModeChngFailed)
- DST411 ($DebugTrapAddr)
- DST412 ($AssertCode)
-
Here are the other values you wanted.
You can see that DST385 does increase as it appears all the PLCs have tripped watchdog another couple times since I put them back in RUN. Seems I forgot to put the Test PLC back into run, so it is still showing 0 for watchdog reboot count.
Thanks,
-
Here are the other values you wanted.
You can see that DST385 does increase as it appears all the PLCs have tripped watchdog another couple times since I put them back in RUN. Seems I forgot to put the Test PLC back into run, so it is still showing 0 for watchdog reboot count.
Thanks,
The value in DebugTrapAddr suggests that this was an OS crash.
-
So looking at System Status I did notice that I had lots of DLWX failed comms on one of the PLCs. Would these cause the OS to crash?
-
So looking at System Status I did notice that I had lots of DLWX failed comms on one of the PLCs. Would these cause the OS to crash?
Honestly nothing should ever cause it to crash, but failure conditions are always the harder thing to test.
I don't have a map file from the exact build you're running, but that address does point to some network stuff in some builds. The only way to be certain what's happening is to give you a new build of the OS with the map updated, then we can know where in the code it was when it went on walkabout. That doesn't always tell us what we need to know, but it's a start.
Not to shirk responsibility, but the ever changing network environment makes network issues a moving target. The issue we had last year where TCP servers (generally Modbus) were locking up ended up being a bug that I'm 90% sure had been in the stack since day 1, but changing network hardware (radios in this case) exposed it 10 years after initial Do-more release. Since there is always something new happening, sometimes thing just pop up. All you can do is deal with it when it happens, but it makes us scratch our heads at times.
-
Ok, on one of the PLCs I patched this code to stop the DLWX comms to see if that stops the watchdogs.
I am willing to try a new OS build to help figure this out as well.
-
I also sat watching the Status for a while and noticed this popping up on occasion. Something to worry about, or possibly related? Seems to correspond to some ping instructions failing.
-
I also sat watching the Status for a while and noticed this popping up on occasion. Something to worry about, or possibly related? Seems to correspond to some ping instructions failing.
Possibly related. That is the TCP stack getting its panties in a bunch, generally due to something not working as expected.
-
Ok, leaning more toward the panic thing being the issue. After stopping the DLWX errors I still got 3 more hardware watchdogs in roughly 24 hours. What do you need me to do in order to help narrow this issue down?
-
Ok, leaning more toward the panic thing being the issue. After stopping the DLWX errors I still got 3 more hardware watchdogs in roughly 24 hours. What do you need me to do in order to help narrow this issue down?
Panics only happen when the stack gets something it isn't sure how to deal with. That suggests some unusual network traffic.
I'm eyeballs deep in some stuff right now, but let me see if I can get a beta build together with an up-to-date map file. That allows me to translate the panic address to the source function that is complaining. Once we know that, we'll figure out what's next. It generally involves peeling the onion with debug builds that dump diagnostic info.
-
So, messing with this some more.
At the locations where these BRX PLCs are having the issues, only the PLC running this specific program is having the watchdog timeouts, while 3 other BRX PLCs on the same network are not having this issue.
On my Test BRX I stopped running any programs that performed any comms actions and I am still getting the hardware watchdogs. On this same network I loaded this same program to a BRX without the I/O modules and it hasn't had the watchdog timeouts. I am now installing the I/O to see if this PLC starts getting watchdog timeouts. I will let you know.
-
If you can reliably duplicate it on a test system, we just need to dupe it here and fix it.
-
Duplicated my answer here since it was pertinent to this thread as well.
So turns out I had a bug in the program for the last 2 years that caused the SETUPIP to execute once a minute. It appears that after enough of these it causes the OS to crash, my guess based on your comments is that occasionally I am unlucky and the crash happens during the flash writing operation and the PW config is lost. At least that is my assumption.
I tested this by having the instruction fire once per second and I very consistently crash the OS after 1408 seconds. Each time it did this I also lost the password configuration. Figured this might be useful for you to see what exactly is happening to cause the OS to crash after writing the IP stuff 1408 times. Related thread: https://forum.hosteng.com/index.php?topic=3852.0
What is interesting is that it doesn't affect the H2 Do-more this same way as those are operating in the field without a hitch even saddled with this bug. That or they are on older firmware(many on 2.7.5) that may not be affected for some reason.