Two networks learn; an independent evaluator decides if they improved
play inside the
environment
copies of self
as opponents
trained weights
policy net
as opponents
experience
uniform replay
recent transitions
reservoir
all-time sample
best-response net
chases the highest-value move
average-policy net
how it tends to play over time
two networks: Neural Fictitious Self-Play
independent evaluator
frozen agent vs fixed baselines, not the reward curve
win rate
reported with 95% confidence intervals