ml4t-rl-execution · git:20260901.c415df0 · 2026-09-01 · sha256 c3920babfac6f2ed
ml4t-rl-execution git:20260901.c415df0A
Immutable. This exact content is served forever at /api/v1/blob/c3920babfac6f2ed.
---
name: ml4t-rl-execution
description: "Reinforcement learning for trade execution and hedging. Use when optimizing execution algorithms or dynamic hedging policies."
when_to_use: "Use when building adaptive execution strategies that minimize market impact beyond static TWAP/VWAP schedules"
dependencies: [run-backtest, cost-model]
metadata:
book_chapters: "21"
library: "ml4t-backtest"
paths: ["**/*backtest*.py", "**/*strategy*.py", "**/*engine*.py", "**/*execution*.py", "**/*rl*.py", "**/*impact*.py"]
---
# RL for Trade Execution
Fixed execution schedules (TWAP, VWAP) ignore real-time market conditions. An RL agent adapts its execution rate based on order book state, reducing implementation shortfall.
## The Problem
Executing a large order at a fixed rate creates predictable market impact. A 100K-share TWAP sell ignores favorable liquidity bursts and pushes through thin books. The result: 20-50 bps of avoidable shortfall on institutional orders, compounding across thousands of trades per year.
## The Pattern
Model execution as a finite-horizon MDP. State: remaining shares, time left, volume, spread. Action: execution rate. Reward: negative implementation shortfall.
### WRONG
```python
# Static TWAP - ignores market conditions entirely
def twap_execute(total_shares: int, n_slices: int) -> list[int]:
base = total_shares // n_slices
remainder = total_shares % n_slices
return [base + (1 if i < remainder else 0) for i in range(n_slices)]
schedule = twap_execute(100_000, 20) # Same size every slice, blind to liquidity
```
### CORRECT
```python
import gymnasium as gym
import numpy as np
from gymnasium import spaces
class ExecutionEnv(gym.Env):
"""Agent decides what fraction of remaining shares to execute each step."""
def __init__(self, total_shares=100_000, n_steps=20):
super().__init__()
self.total_shares, self.n_steps = total_shares, n_steps
self.observation_space = spaces.Box(0, 1, shape=(4,), dtype=np.float32)
self.action_space = spaces.Box(0, 1, shape=(1,), dtype=np.float32)
def reset(self, seed=None, options=None):
super().reset(seed=seed)
self.remaining, self.step_idx = self.total_shares, 0
self.arrival_price = 100.0
return self._obs(), {}
def step(self, action):
shares = int(np.clip(action[0], 0, 1) * self.remaining)
impact = 0.0001 * (shares / 5000) # Linear market impact
exec_price = self.arrival_price * (1 + impact)
shortfall = (exec_price - self.arrival_price) / self.arrival_price
reward = -abs(shortfall) * shares / self.total_shares
self.remaining -= shares
self.step_idx += 1
done = self.step_idx >= self.n_steps or self.remaining <= 0
if done and self.remaining > 0:
reward -= 0.01 # Non-completion penalty
return self._obs(), reward, done, False, {}
def _obs(self):
return np.array([
self.remaining / self.total_shares, self.step_idx / self.n_steps,
np.random.uniform(0.01, 0.05), # spread
np.random.uniform(0.3, 1.0), # volume ratio
], dtype=np.float32)
```
## Guardrails
- **Non-completion penalty is mandatory** - without it the agent learns zero-trade is optimal
- **Normalize all state features** - raw share counts and prices break learning
- **Validate against TWAP baseline** - if RL underperforms TWAP, the environment is misconfigured
- **Use square-root impact for large orders** - linear impact underestimates cost at scale
- **Episode = one parent order** - do not mix multiple orders into one episode
## Production Implementation
`ml4t-backtest` provides execution simulation with realistic market impact:
```python
from ml4t.backtest import BacktestConfig, CommissionType, Engine
from ml4t.backtest.config import SlippageType
from ml4t.backtest.execution.impact import SquareRootImpact
from ml4t.backtest.execution.limits import VolumeParticipationLimit
config = BacktestConfig(
commission_type=CommissionType.PER_SHARE,
commission_per_share=0.005,
slippage_type=SlippageType.VOLUME_BASED,
slippage_rate=0.001,
)
env_engine = Engine(
feed,
strategy,
config,
market_impact_model=SquareRootImpact(volatility=0.02),
execution_limits=VolumeParticipationLimit(max_participation=0.05),
)
```
## Checklist
- [ ] Environment has both time pressure and execution cost in the reward
- [ ] State is normalized (fractions, ratios) not raw values
- [ ] Non-completion is penalized (agent must finish the order)
- [ ] Trained agent beats TWAP baseline on test episodes
- [ ] Action space bounded (cannot execute more than remaining shares)