In the world, most of the successes are results of long-term efforts. The reward of success is extremely high, but be-fore that, a long-term investment process is required. People who are "myopic" only value short-term rewards and are unwill-ing to make early-stage investments, so they hardly get the ulti-mate success and the corresponding high rewards. Similarly, for a reinforcement learning (RL) model with long-delay rewards, the discount rate determines the strength of agent's "farsightedness". In order to enable the trained agent to make a chain of correct choices and succeed finally, the feasible region of the discount rate is obtained through mathematical derivation in this paper firstly. It satisfies the "farsightedness" requirement of agent. Af-terwards, in order to avoid the complicated problem of solving implicit equations in the process of choosing feasible solutions, a simple method is explored and verified by theoreti cal demon-stration and mathematical experiments. Then, a series of RL ex-periments are designed and implemented to verify the validity of theory. Finally, the model is extended from the finite process to the infinite process. The validity of the extended model is veri-fied by theories and experiments. The whole research not only reveals the significance of the discount rate, but also provides a theoretical basis as well as a practical method for the choice of discount rate in future researches.