• Some users have recently had their accounts hijacked. It seems that the now defunct EVGA forums might have compromised your password there and seems many are using the same PW here. We would suggest you UPDATE YOUR PASSWORD and TURN ON 2FA for your account here to further secure it. None of the compromised accounts had 2FA turned on.
    Once you have enabled 2FA, your account will be updated soon to show a badge, letting other members know that you use 2FA to protect your account. This should be beneficial for everyone that uses FSFT.

linux smp client hanging

tkam

[H]ard|DCer of the Month - Dec. 2012
Joined
Dec 18, 2007
Messages
436
I remember running into similar problems with the linux smp client hanging before attempting to send the results. But now it seems to be hanging after it sends the results and right before it starts the next WU.

Code:
[14:20:44] Completed 250000 out of 250000 steps  (100%)

Writing checkpoint, step 5750000 at Wed Nov 19 09:20:44 2008

Writing final coordinates.

Average load imbalance: 4.0 %
Part of the total run time spent waiting due to load imbalance: 2.8 %
Steps where the load balancing was limited by -rdd, -rcon and/or -dds: Z 0 %

	Parallel run - timing based on wallclock.
               NODE (s)   Real (s)      (%)
       Time:  65886.000  65886.000    100.0
                       18h18:06
               (Mnbf/s)   (GFlops)   (ns/day)  (hour/ns)
Performance:     54.129      8.251      0.656     36.603

gcq#0: Thanx for Using GROMACS - Have a Nice Day

[14:21:46] 
[14:21:46] Finished Work Unit:
[14:21:47] - Reading up to 21132000 from "work/wudata_01.trr": Read 21132000
[14:21:49] trr file hash check passed.
[14:21:49] - Reading up to 4506556 from "work/wudata_01.xtc": Read 4506556
[14:21:49] xtc file hash check passed.
[14:21:49] edr file hash check passed.
[14:21:49] logfile size: 179755
[14:21:49] Leaving Run
[14:21:50] - Writing 26052775 bytes of core data to disk...
[14:21:53]   ... Done.
[14:22:02] - Shutting down core
[14:22:02] 
[14:22:02] Folding@home Core Shutdown: FINISHED_UNIT
Error encountered before initializing MPICH
[14:25:12] CoreStatus = 64 (100)
[14:25:12] Sending work to server


[14:25:12] + Attempting to send results
[14:26:44] + Results successfully sent
[14:26:44] Thank you for your contribution to Folding@Home.
[14:26:44] + Starting local stats count at 1

As you can see it stops at the "[14:26:44] + Starting local stats count at 1" and it just sits there. This is happening on two of my boxes (one is a ubuntu 8.04 dual opteron box and the other is a ubuntu 8.04 VM box). On each box I've let the client sit at that state for hours and hours and it never gets past there. I'm using the unified 6.02 client, would the 6.23 beta help at all?
 
What core is being used when crunching these work units? Is it the A1 or the A2 core? If it's the A2 core, then welcome to the club. That's why I gave up on everything except for the GPU client. I was sick and tired of babysitting 6 SMP clients and having folding time wasted due to idling boxen. Also, look to see what the project number of the work units are that are hanging as sometimes there are some crappy work units which may do something like this.

Since I stopped the SMP client, I haven't really kept up on anything so I have no clue if the A2 core issue has been fixed or not if that is your problem. Someone else may be able to chime in.

 
I'm getting them hanging in that spot on pretty much every A2 work unit lately. It looked like they had it taken care of for a while, but they're back. :(
 
Looks like the A2 core guess I should have searched some more first I didn't realize this was a known/on-going issue.

It's strange though I have two other boxes running the linux-smp client using the A2 core and they haven't had a single problem. Though those are both intel quad core boxes, maybe that has something to do with it.
 
try the 6.23 beta, it doesn't hurt since you already have problems anyway.

 
Yeah I just switched the the two I'm having issues with over to the 6.23 beta. I'll report back if it helps.
 
I had six clients running the A2 core and every damn one of them would hang after sending back the finished work unit. That said, it wouldn't do it with every work unit that was processed with the A2 core but it was often enough that I got sick of having to fix at least a couple of them a day.

I never had to restart the client as that wasn't the problem. After the work unit was finished and started uploading, all four of the A2 core processes would disappear. However, after the work unit was done uploading, at least one A2 core process would return even before another work unit was grabbed. This caused the client to hang. After killing the A2 core process or processes, the client would continue on as if nothing was wrong. If none of the A2 core processes started up again before another work unit was downloaded, it wouldn't hang.

If Stanford can fix this, I'll think about moving some or all of my quads back over to the SMP client, but I won't even consider it until it's fixed.

 
Always look on the bright side, at that point if you do a ctl+c and hit the up arrow key once you will be at the command to start. Hit enter and it will fold on;)

 
I will chime in here because I was intending to post a thread about this issue but have been procrastinating for almost a week. I have 14 Linux SMP clients running currently. They're all Intel quad based clients except for one. There is no difference in regards to susceptibility when it comes to the hanging issue. It only occurs with A2 WUs. Another thing I noticed is the problem getting worse not better. Whereas the hanging used to afflict roughly 40% of my completed A2 WUs, the problem has worsened and I estimate that 70% of these A2 units hang after sending in their results. Compound this fact with the near elimination of A1 WUs from the repertoire of work that the Linux SMP clients receive from Stanford lately, and you can just see the scope of the problem I'm dealing with. To recap:

I have a total 14 Linux SMP clients but could potentially run more if I fix a downed server.
None of these clients are presently being monitored by either FahSpy or FahMon mandating regular checks.
There is a minimum 70% likelihood they will hang after submitting results.
Almost 100% of the current work being dispatched from Stanford to these clients is A2 core-based that is prone to hanging.
Each of these clients completes their WUs in a day or less (~10 hangs out of 14 completed WUs per day).
See the problem?

 
Weird, I have zero hang on my 4 VM instances in over 1 month... Most is 6.02 excepted 1 which run 6.23 to fix a unrelated issue.

 
Weird, I have zero hang on my 4 VM instances in over 1 month... Most is 6.02 excepted 1 which run 6.23 to fix a unrelated issue.
What does 6.23 fix? I have too many clients to update and unless I can get some confirmation that it will help with this issue, I won't install the new version.
 
It improve the handling of EUE, not much more details beside this.

 
I just upgraded to 6.23 beta and the first A2 unit out of the gate finished without the hang, so maybe I have a bit less babysitting in my future. :D


2.png
 
I just upgraded to 6.23 beta and the first A2 unit out of the gate finished without the hang, so maybe I have a bit less babysitting in my future. :D
How did you upgrade to the new client? Is it a simple .exe replacement?
 
I've upgraded both of my problem boxes to the 6.23 client and so far no issues, one of the boxes has finished two wu's and the other has finished one.

 
Yea, its a replacement fah6, just drop it in.

Edit: here's a link, its not on the main download page
I tried the updated exe on one of my clients and now I can't execute the client. I keep on receiving the 'permission denied' message whenever I try to launch it. :confused:
 
A what? No, I didn't. What does this command do? Grant administrative rights?

It makes the file executable. If you're overwriting the file, it should keep the attributes of the old one, but sometimes you have to reset it.
 
It makes the file executable. If you're overwriting the file, it should keep the attributes of the old one, but sometimes you have to reset it.
I see. I didn't know we have to do this in Linux. Nothing was necessary with the original file to make it executable. If I delete the original exe file before downloading the new one, will it execute without the command?
 
Interesting. I've ALWAYS had to do a chmod on the executable, except when overwriting one. But then again, I've learned a long time ago that all the F@H clients just hate me.


2.png
 
I've always had to chmod it as well. But to report back again on 6.23 beta both of my problem boxes have been running trouble free since the switch to 6.23.



 
I forgot to add, dropping in 6.23beta has made it considerably better, but I've still a couple hangs. I'm going to try wiping the directory and reinstalling, but I'm too lazy right now.


2.png
 
I see. I didn't know we have to do this in Linux. Nothing was necessary with the original file to make it executable. If I delete the original exe file before downloading the new one, will it execute without the command?

Interesting. I've ALWAYS had to do a chmod on the executable, except when overwriting one. But then again, I've learned a long time ago that all the F@H clients just hate me.


2.png

I've always had to chmod it as well. But to report back again on 6.23 beta both of my problem boxes have been running trouble free since the switch to 6.23.




Much of it depends on how you extract the file. With command line based extractions I have found that I have to chmod it to make it executable. In the cases where I have used Ark to extract the file, I did not need to change anything to make it executable.

 
Back
Top