Reinforcement Learning for Pricing American Options
En cours de chargement...
Date
Authors
Nom de la revue
ISSN de la revue
Titre du volume
Éditeur
Université d'Ottawa | University of Ottawa
Résumé
This thesis develops a martingale-based reinforcement-learning (RL) framework for pricing European and American options. For each learning mode, a shared European option price surface is learned from risk-neutral paths and then used as the baseline for an American early-exercise correction. Offline learning uses complete simulated paths, whereas online learning uses one-step martingale increments. The American option premium-learning
stage uses a penalized, entropy-regularized two-action stopping formulation, and prices are evaluated through the deterministic stopping rule obtained after training. Under Black–Scholes dynamics, maximum American execution-price errors are 2.40% offline and 2.97% online. Under Merton jump–diffusion dynamics, a direct transfer of the Black–Scholes implementation did not adequately capture early exercise. Improving state-space and
exercise-date coverage and reorganizing the updates while retaining the two-stage formulation yields offline and online mean errors of 1.31% and 1.09%, with maxima of 2.28% and 2.09%, respectively. Both methods recover meaningful early-exercise behaviour.
Description
Mots-clés
Pricing Options, American Options, Reinforcement Learning
