You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Reference simulator and stored results for Lyapunov-guided safe RL with risk-budget feasibility in energy-aware O-RAN scheduling: CVaR tail constraint, LCB safety filter, PID dual, and a swept classical drift-plus-penalty frontier.
Does richer tail-risk feedback help a language model write a better trading reward? A pre-registered study across 11 models, five feedback arms and 568 seeds per comparison unit. MSc dissertation, UCL Institute of Finance and Technology.
A Reinforcement Learning MVP (Minimum Viable Product) for Condition-Based Maintenance (CBM) using industrial equipment temperature sensor data. This project implements a sophisticated QR-DQN (Quantile Regression Deep Q-Network) agent to learn optimal maintenance policies balancing risk mitigation and cost minimization.
Code for UCB-BQRL, a model-based optimistic reinforcement learning algorithm for risk-sensitive control with smoothed quantile objectives under unknown transition dynamics.
Adaptive risk-aware reinforcement learning framework for time-constrained navigation tasks. This MSc thesis studies how an RL agent can adapt its risk attitude when mission time changes, using a hierarchical controller that switches from a risk-aware CVaR-constrained policy to a risk-seeking policy when time pressure makes switching benificial.