Host Engineering Forum

General Category => Do-more CPUs and Do-more Designer Software => Topic started by: Scot on March 01, 2019, 12:02:45 PM

Title: Hardware watchdog timeout
Post by: Scot on March 01, 2019, 12:02:45 PM
I just had my first major glitch with a H2-DM1E.
In the event log it shows a Hardware Watchdog Timeout. This caused a reset and a loss of all variables - but the program was still intact. So all tuning parameters, including for PID loops were gone.

Is there more troubleshooting I can do to narrow down what caused this?
Title: Re: Hardware watchdog timeout
Post by: BobO on March 01, 2019, 04:48:27 PM
If the system hasn't been restarted more than 4 times since the reboot, we may be able to glean some information from some postmortem entries in DST memory locations 400-409. You can pull those up in a data view.

That assumes that the retentive memory is intact, which when you say "loss of all variables", I'm guessing the retentive memory was cleared?

One thing you should always do is to store a copy of the retentive memory. You can do that with the Retentive Memory Manager. That makes it very quick to restore the PLC in the event of a hardware failure.
Title: Re: Hardware watchdog timeout
Post by: Scot on March 04, 2019, 01:31:49 PM
Thanks for the reply BobO.

It hasn't been rebooted at all since it self-restarted. The DST400-409 values are:
DST400=806
DST401=1075183616
DST402-409 = 0

Yes, all the retentive memory was cleared. Including a range of 5000 signed DWords, 5000 unsigned Words, 5000 bytes for Die serial number tracking and 15000 signed DWords, 15000 unsigned Words, and 15000 bytes for Roll serial number tracking.

I'll definitely try out the Memory Image Manager but I'll wait until we are shut down, just in case.

I can also add some PIDINIT commands but the problem with those is I need to remember to update them whenever loop tuning is changed.
Title: Re: Hardware watchdog timeout
Post by: BobO on March 04, 2019, 01:44:31 PM
The relevant values would have been DST402-403. I'm not sure we would have learned anything though, since those generally contain mode change and initialization information. There are other locations that also could have been of use, but they are no doubt cleared as well.

Not super thrilled about the retentive memory getting wiped. Unless the battery was bad, it should be very difficult for that to happen. Even less thrilled about the reboot. Some questions:
1. How old is the unit?
2. Anything else going on that could've created more noise than usual?
3. Changed network load/device/conditions?
4. When was the last time the PLC was rebooted?
5. Recent changes to PLC program?

We can dig further and will do whatever you wish to try to solve it.
Title: Re: Hardware watchdog timeout
Post by: Scot on March 06, 2019, 12:20:11 PM
It's one of the first Do-More's I installed. When the Do-More Designer first added the ability to translate from a DirectSoft program was when I made the conversion. I had submitted some bug reports at the time on the timer conversion scale being off. I would guess no older than 2014. There was no power fluctuation at the time.

One user was looking at Roll history at the time of the crash which includes a search task that has a For-Next loop. It's been in use for almost a year with no problems though.

The ECOM100 is on an isolated network for all of the PLC's and C-Mores. The expansion bases for the PLC are connected to the port on the Do-More on a different isolated network. Each Do-More and it's expansions has it's own network.

It's almost never rebooted. The last time it changed to program mode was when I added some PID loops about 3 months ago during a shutdown. Most of the time it just gets Run mode edits.

There was no changes being made at the time it crashed. The Designer wasn't open so no PC was connected at the time.

My best guess was a glitch in the serial number tracking that caused a recursive loop, but it hasn't had problems before and I thought the task was set to yield every 100 μs. The guy looking at the Roll list wasn't doing a search though, just viewing the list. So there should have been no loop. Maybe there was a reference the a variable out of range. The code uses references like RollSer[V8013]. V8013 couldn't be negative but maybe there was a glitch that made it larger than the defined RollSer range (0 to 14999). I don't know what would happen in that case.

It's not complicated code and has never bothered before though. I could send it to you if you want to see it.
Title: Re: Hardware watchdog timeout
Post by: BobO on March 06, 2019, 12:32:44 PM
We like to be proactive on these things. I don't hear anything there that would give me pause, but it rebooted for some reason. Please let us know if it happens again.
Title: Re: Hardware watchdog timeout
Post by: Scot on March 06, 2019, 12:42:08 PM
I've generated new images of the memory using the contents from the PLCs on 2 of the systems (we have 6 Do-Mores in operation).

When I did that, I also noticed a built in memory block of strings called LastERR. There's 8 strings and the errors are all 2 different errors that are repeated and they do appear to be from the Die and Roll serial number lists (Dies range from 0 to 4999 and Rolls range from 0 to 14999). The 2 errors are "DWordArray index out of range: 5403 > 4999" and "DWordArray index out of range: 32992 > 14999". So maybe I do have a bug somewhere.

Edit: 7 in operation. I forgot the truck dumps.
Title: Re: Hardware watchdog timeout
Post by: BobO on March 06, 2019, 12:47:38 PM
I've generated new images of the memory using the contents from the PLCs on 2 of the systems (we have 6 Do-Mores in operation).

When I did that, I also noticed a built in memory block of strings called LastERR. There's 8 strings and the errors are all 2 different errors that are repeated and they do appear to be from the Die and Roll serial number lists (Dies range from 0 to 4999 and Rolls range from 0 to 14999). The 2 errors are "DWordArray index out of range: 5403 > 4999" and "DWordArray index out of range: 32992 > 14999". So maybe I do have a bug somewhere.

Those are the OS preventing array references from escaping the memory block.

Shouldn't be possible to crash the CPU with bad data. Doesn't mean you can't.