Blog
DRL Agent Deployment: What Breaks Between Training and Live
Most DRL agent deployment failures are not model failures — they are feature parity failures, and they produce no error message. The DRL agent described here converged cleanly: reward curves…
What Goes in a DRL Trading Agent’s Observation Vector
A PPO agent trained for 2 million steps on EURUSD 15-minute data showed flat reward curves across all 5 seeds. Switching to SAC changed nothing. Adjusting the learning rate changed…
Which DRL Algorithm for FX Trading — and Why the Max Drawdown Number Tells You Everything
DQN produced a higher annualized return than PPO in a commodity futures benchmark — 47.6% against 15.7%. It also produced a maximum drawdown of -16.6%, against PPO’s -0.75%. The second…