Machine learning has shown potential to improve portfolio optimization but whether that can translate into practical real-world applications remains uncertain. Recently models have started combining machine learning with classic theory and/or other machine learning models, chasing higher performance and more transparency. Due to the novelty of these models the practical extent of what they can contribute to portfolio management remains unknown.
In this thesis we built on a recently proposed idea and constructed a deep reinforcement learning model with a transformer neural network integrated with the Black-Litterman framework for portfolio optimization. The model was trained on 14 years of daily stock returns from the FTSE 100 index; a much broader and longer data set compared to that used in the model’s introductions. The model learned dynamic views of the market and of its own risk preference so it could adapt to changing market regimes.
When using back testing methods our model outperformed traditional benchmark strategies in both cumulative returns and Sharpe ratios. Despite strong results in this thesis, and in recent literature, there are inherent limitations to machine learning that stillstand in the way of real-world viability.