Two networks learn; an independent evaluator decides if they improved

play inside the environment copies of self
as opponents
experience
uniform replay recent transitions
reservoir all-time sample
best-response net chases the highest-value move
average-policy net how it tends to play over time
two networks: Neural Fictitious Self-Play
trained weights policy net
independent evaluator frozen agent vs fixed baselines, not the reward curve win rate reported with 95% confidence intervals