Reinforcement Learning Beyond Greedy Optimisation for Accelerator Control with Delayed Consequences JACoW Publishing