Destination -> Like, close to the destination? I don't see how that's hard.
Happy -> you can use customer feedback for this
Folded -> this is indeed the trickiest one, but I think well within the capabilities of modern vision models.
Really? What about fires? Falling off cliffs? Causing others to crash?
Your "examples" are all hand-wavy and vague and no good to train an RL agent. You've also not provided a reward function.
Destination -> Like, close to the destination? I don't see how that's hard.
Happy -> you can use customer feedback for this
Folded -> this is indeed the trickiest one, but I think well within the capabilities of modern vision models.