Sorry, but so far the discussion above is pure speculation. If your users cannot login - the server most likely "crashed", but there are millions of reasons why this can happen. It can be Aware IM fault, but it can also be the application's fault or an incorrect system setup.
So whenever you have a problem like this you need to run Aware IM as a service and keep the log file (wrapper.log). As soon as you notice the problem you should capture the time of the problem and save the log away immediately, so that it can be analysed around the time of the problem. If you cannot interpret the log yourself, send it to us (and pay for support). We should then take it from there.
About hanging connections. This should never ever happen. If it does it's a serious bug. But you need to prove that this is indeed what's happening. Can someone prepare a test when you definitely prove that connections are left hanging? (this is very unlikely or it happens under some very esoteric circumstances, otherwise any system wouldn't last very long)
About execution contexts. If a process fails it should be removed from the execution context table. If it's not, it's a serious bug. Can someone prepare a test that proves that this is a bug?
The system should timeout processes in the execution_context table automatically, so it should never grow too big. If it doesn't, then it's a serious bug again. Can someone prove that this is indeed a bug?
We have checked the code that deals with connections and execution_contexts multiple times and we don't see anything wrong with it. That said, there can be scenarios (in theory) that haven't been checked, but we need your help in identifying these scenarios if you are sure that there is indeed a bug in the system.