Host Engineering Forum

General Category => Do-more CPUs and Do-more Designer Software => Topic started by: JeffS on September 29, 2023, 05:50:26 PM

Title: BRX Hardware Watchdog Timeout
Post by: JeffS on September 29, 2023, 05:50:26 PM
I have a few PLCs that are pretty consistently triggering the hardware watchdog timeout.  Attached are the system logs for the two PLCs.  Is there anything I can do to narrow down what could be causing this?   For the Test BRX Maestro it has nothing wired up to any of the cards attached to the BRX PLC.  On my test BRX I plan to unplug the connectors on every card so they don't have external power, just to see if that affects the hardware timeout occurrence.

Carnegie PLC: 2.9.6 OS, 1.1.0 Booter

Test PLC: 2.9.7 OS, 1.1.0 Booter

What can I do to help narrow this down to a particular card or whatever is actually causing the timeout?

Thanks,
Jeff
Title: Re: BRX Hardware Watchdog Timeout
Post by: Greg on October 02, 2023, 12:49:37 PM
Normally, a hardware watchdog trigger is caused by something physical like:
The System Data Word, DST385 ($WatchdogReboots) shows a count of how many times it has rebooted.
DST400-409 contain codes for us (Host Engineering) that may help narrow down what is happening. Can you post those?
Title: Re: BRX Hardware Watchdog Timeout
Post by: JeffS on October 04, 2023, 11:07:35 AM
Attached are some screen shots of the contents of DST385 and DST400-409.  I have this happening on another PLC in the field as well at a separate location.
Where DST385 is 0 is because the PLC had dropped out of run after reaching 10, which each of these PLCs had all done at the time I logged in to get the data.
Title: Re: BRX Hardware Watchdog Timeout
Post by: Greg on October 05, 2023, 10:35:36 AM
Attached are some screen shots of the contents of DST385 and DST400-409.  I have this happening on another PLC in the field as well at a separate location.
Where DST385 is 0 is because the PLC had dropped out of run after reaching 10, which each of these PLCs had all done at the time I logged in to get the data.
The codes in DST400-409 look normal. But DST385 ($WatchdogReboots) = 0 is a bit troublesome.

Does this ever increment?

Also, since these all look normal (except for the DST385), what are these values?
- DST410 ($ModeChngFailed)
- DST411 ($DebugTrapAddr)
- DST412 ($AssertCode)
Title: Re: BRX Hardware Watchdog Timeout
Post by: JeffS on October 05, 2023, 12:55:45 PM
Here are the other values you wanted.

You can see that DST385 does increase as it appears all the PLCs have tripped watchdog another couple times since I put them back in RUN. Seems I forgot to put the Test PLC back into run, so it is still showing 0 for watchdog reboot count.

Thanks,
Title: Re: BRX Hardware Watchdog Timeout
Post by: BobO on October 05, 2023, 03:27:57 PM
Here are the other values you wanted.

You can see that DST385 does increase as it appears all the PLCs have tripped watchdog another couple times since I put them back in RUN. Seems I forgot to put the Test PLC back into run, so it is still showing 0 for watchdog reboot count.

Thanks,

The value in DebugTrapAddr suggests that this was an OS crash.
Title: Re: BRX Hardware Watchdog Timeout
Post by: JeffS on October 05, 2023, 04:50:15 PM
So looking at System Status I did notice that I had lots of DLWX failed comms on one of the PLCs.  Would these cause the OS to crash?
Title: Re: BRX Hardware Watchdog Timeout
Post by: BobO on October 05, 2023, 05:03:48 PM
So looking at System Status I did notice that I had lots of DLWX failed comms on one of the PLCs.  Would these cause the OS to crash?

Honestly nothing should ever cause it to crash, but failure conditions are always the harder thing to test.

I don't have a map file from the exact build you're running, but that address does point to some network stuff in some builds. The only way to be certain what's happening is to give you a new build of the OS with the map updated, then we can know where in the code it was when it went on walkabout. That doesn't always tell us what we need to know, but it's a start.

Not to shirk responsibility, but the ever changing network environment makes network issues a moving target. The issue we had last year where TCP servers (generally Modbus) were locking up ended up being a bug that I'm 90% sure had been in the stack since day 1, but changing network hardware (radios in this case) exposed it 10 years after initial Do-more release. Since there is always something new happening, sometimes thing just pop up. All you can do is deal with it when it happens, but it makes us scratch our heads at times.
Title: Re: BRX Hardware Watchdog Timeout
Post by: JeffS on October 05, 2023, 05:13:23 PM
Ok, on one of the PLCs I patched this code to stop the DLWX comms to see if that stops the watchdogs. 

I am willing to try a new OS build to help figure this out as well.
Title: Re: BRX Hardware Watchdog Timeout
Post by: JeffS on October 05, 2023, 06:16:47 PM
I also sat watching the Status for a while and noticed this popping up on occasion.  Something to worry about, or possibly related?  Seems to correspond to some ping instructions failing.
Title: Re: BRX Hardware Watchdog Timeout
Post by: BobO on October 05, 2023, 07:33:53 PM
I also sat watching the Status for a while and noticed this popping up on occasion.  Something to worry about, or possibly related?  Seems to correspond to some ping instructions failing.

Possibly related. That is the TCP stack getting its panties in a bunch, generally due to something not working as expected.
Title: Re: BRX Hardware Watchdog Timeout
Post by: JeffS on October 06, 2023, 05:53:45 PM
Ok, leaning more toward the panic thing being the issue.  After stopping the DLWX errors I still got 3 more hardware watchdogs in roughly 24 hours.  What do you need me to do in order to help narrow this issue down?
Title: Re: BRX Hardware Watchdog Timeout
Post by: BobO on October 09, 2023, 11:37:24 AM
Ok, leaning more toward the panic thing being the issue.  After stopping the DLWX errors I still got 3 more hardware watchdogs in roughly 24 hours.  What do you need me to do in order to help narrow this issue down?

Panics only happen when the stack gets something it isn't sure how to deal with. That suggests some unusual network traffic.

I'm eyeballs deep in some stuff right now, but let me see if I can get a beta build together with an up-to-date map file. That allows me to translate the panic address to the source function that is complaining. Once we know that, we'll figure out what's next. It generally involves peeling the onion with debug builds that dump diagnostic info.
Title: Re: BRX Hardware Watchdog Timeout
Post by: JeffS on October 10, 2023, 12:19:49 PM
So, messing with this some more. 

At the locations where these BRX PLCs are having the issues, only the PLC running this specific program is having the watchdog timeouts, while 3 other BRX PLCs on the same network are not having this issue.

On my Test BRX I stopped running any programs that performed any comms actions and I am still getting the hardware watchdogs.  On this same network I loaded this same program to a BRX without the I/O modules and it hasn't had the watchdog timeouts.  I am now installing the I/O to see if this PLC starts getting watchdog timeouts.  I will let you know.



Title: Re: BRX Hardware Watchdog Timeout
Post by: BobO on October 11, 2023, 11:08:06 AM
If you can reliably duplicate it on a test system, we just need to dupe it here and fix it.
Title: Re: BRX Hardware Watchdog Timeout
Post by: JeffS on October 12, 2023, 07:32:00 PM
Duplicated my answer here since it was pertinent to this thread as well.

So turns out I had a bug in the program for the last 2 years that caused the SETUPIP to execute once a minute.  It appears that after enough of these it causes the OS to crash, my guess based on your comments is that occasionally I am unlucky and the crash happens during the flash writing operation and the PW config is lost.  At least that is my assumption.   

I tested this by having the instruction fire once per second and I very consistently crash the OS after 1408 seconds.  Each time it did this I also lost the password configuration.  Figured this might be useful for you to see what exactly is happening to cause the OS to crash after writing the IP stuff 1408 times.   Related thread: https://forum.hosteng.com/index.php?topic=3852.0

What is interesting is that it doesn't affect the H2 Do-more this same way as those are operating in the field without a hitch even saddled with this bug.  That or they are on older firmware(many on 2.7.5) that may not be affected for some reason.