The server ran out of memory in the middle of an order
A small cloud server ran out of memory at the busiest minute of the day. The operating system killed the bot halfway through placing an order, and the restart nearly placed it again.
What went wrong
Our bot runs on a small cloud server with well under 1 GB of memory. One morning, a few minutes after the open, several strategies fired in the same minute while a background job was also recording market data. Memory ran out, and the operating system’s out-of-memory killer stopped the bot instantly, with no error message and no chance to clean up.
It happened in the middle of placing a multi-leg order. Some legs had been sent, others hadn’t. A watchdog restarted the bot about 30 seconds later, and the fresh process, seeing no record of a completed entry, started to enter again.
What it cost, or could have cost
A half-built position for about half a minute, then the risk of a duplicate one on top of it. On a short-options position, an unplanned doubling of size is exactly the kind of mistake that turns an ordinary bad day into a very bad one.
Root cause
Three things lined up:
- Too little memory headroom. The bot alone used about a third of the server’s memory; at the morning peak, everything together needed more than the server had, and there was no swap space to absorb it.
- The busiest minute is predictable. Most of our strategies act just after the open, so memory peaks at the same time every day: the worst possible moment to crash.
- Restart without memory of intent. The bot recorded an entry only after it was complete. A crash between “sent the first leg” and “finished” left no trace, so the restart assumed nothing had happened.
The fix
- Swap space as a cushion. A 2 GB swap file, with the kernel told to use it only under real pressure (a low “swappiness” setting). It isn’t a substitute for enough memory, but it turns a sudden kill into a brief slowdown.
- Write the intent before acting. Before sending any order, the bot writes “entry attempted” to disk. On restart it checks that flag and the broker’s actual positions before doing anything, so it completes or reconciles a half-built position instead of starting a new one.
- Know how to spot it. An out-of-memory kill leaves no error in the bot’s own log: the process just vanishes and a fresh start appears. The evidence is in the system log:
sudo dmesg | grep -i "killed process"
journalctl -k | grep -i oom
- Reduce the peak (still on our list). Several strategies each load the same large reference files; loading them once and sharing them would cut the morning peak. Swap buys time; it doesn’t remove the cause.
Checklist items
- Your bot restarts automatically after a crash: a watchdog or service manager brings it back within seconds.
- Your bot can be killed at any instant, including mid-order, and restart safely: it records what it is about to do before doing it, and checks the broker’s positions before acting again.
- You monitor the server’s memory use regularly, and get an alert well before it runs out, especially around the busiest minutes of the day.
Was this useful?
More to say, or spotted a mistake? Write to [email protected]. We read everything.