Neuroevolution with the Genetic Algorithm in Python
June 30 2026 - 14 Minutes - Source
pygame - ai - python
this wasn't meant to be the project for this month, but due to the planned one not working out, and me getting the idea for this project, this became this month's project
Here you can see the agents playing one of the two games included! in this case maze!
Hello everyone! I apologize for the late release of this blog post, but the original project I had planned, a Rimworld-clone but with Micheal because I was curious as how the storytellers worked, ended up not working sadly, but luckily I had this idea just a few days ago, and was able to whip it up pretty quick!
this is the third ga project ive done already i need to stop
Introduction
This is the third in my series of genetic algorithm python program, but this one is different, as I am evolving the weights of neural networks, not just a static list of movements.
The scoring, and evolution method are also different for this project vs the last ones.
Technical Details
So lets talk about how this one is different, I originally started with only the neural networks and new scoring, but during that I realized that reproduction changes would help.
The Neural Networks
The networks are very simple, with 7 inputs, two hidden layers of four neurons, and four output neurons, The diagram below shows the full layout, with labels.
The weights are represented very simply as an array in order, its very simple, but worked very well, the direct comment from my code can be seen below!
# 28 to hidden1
# 16 to hidden2
# 16 to output
# array diagram
# [7 h1] [7 h2] [7 h3] [7 h4]
# [4 h1] [4 h2] [4 h3] [4 h4]
# [4 o1] [4 o2] [4 o3] [4 o4]
I chose to do it this way, just due to the fact that the network was so simple, and anything more complex would have been annoying to deal with, and over complicated.
This came back to bite me, as it did make it more difficult to add more neurons, input and hidden, but was still very workable.
The inputs are:
- The type of the tile above it
- The type of the tile below it
- The type of the tile to the left of it
- The type of the tile to the right of it
- The current iteration/step of the round
- The agent's current score
- And a bias node, which is always 1
The tile types are either 1, 0, -1, -2, which means an apple, nothing, a wall, and a collected apple respectively.
Inspecting the weights of the models after the fact, the iteration and score inputs are used, but I am not sure how they are being used.
Why not MNNlibV2?
If you have read some of my earlier blog posts, you may remember MNNlibV2, the Minejerik Neural Network library version 2, and you may be wondering, why didn't I use that, instead of reinventing the wheel?
The simple answer is, MNNlib was not built for situations like this, where training isn't done using gradient descent, and it would've made it much more difficult to edit/evolve and just generally access the neurons.
I could've went back and edited MNNlib to add these features, namely lower level access and different training types, but this would've taken me longer than just reimplementing the basic features I needed.
Also, generally, I have gotten better at programming in the 3 years since I wrote MNNlib, and I did not feel like touching that old of code.
mnnlib is nearly 3 years old oh my god felt like yesterday
Another issue is the very high-level access that MNNlib provides, as it was meant to be very easy, low-friction to use, and I would require direct access to the weight objects to modify them, and generally, the benefits provided by MNNlib allowing me to run the model without writing any code, would not have been worth it for the challenges.
The Scoring
As the development of this project went on, I realized that I would need to modify how scoring works, meaning that scoring ended up quite complex.
Originally, the agents had the ability to not move at all, I removed this later due to the agents learning to just stay still instead of collecting apples, but to prevent this, every time the agent stays still, a 0.09 point penalty is applied to its score.
Once the stay still option was removed, this was not necessary, but the code still exists.
If an agent runs into a wall, or attempts to move into a wall at least, 0.09 points are removed as a penalty, agents still ended up running into walls for some reason, despite being able to see that a wall exists there, so I am not sure.
Also to promote movement, whenever an agent goes to a brand new cell that it hasn't gone before, it gets a modest reward of 0.02 score, but if it goes back to a tile that it has already explored, it gets a small penalty of 0.01 score, to try and prevent agents from quickly going back and forth between two cells, again, despite my best efforts, the agents still ended up doing this, less than without this penalty, but still alternating between cells.
Finally, whenever an agent reaches an apple, it gets 2.5 score added as a reward, once an agent ends up getting an apple for the first time, after a few generations, nearly every agent ends up searching for them.
I still ended up with an issue however, where sometimes, an agent obtains an apple, then ends up receiving many penalties from running into walls, or going over already explored tiles, so the score is put through this formula before final selection takes place, to fully incentivize obtaining apples over everything.
score = (score*((apple_count/5)+1)) + apple_count
I first only had the multiplier, but I still had the issue, so I added the addition of the apple count, this helped eliminate the majority of issues of agents being removed despite collecting a high amount of apples, it still happens in some cases, but is generally better.
The Reproduction System
In the first two genetic algorithm projects, reproduction was a very simple procedure, The best agent is found, the rest are discarded, the best agent is duplicated as many times as needed to fill the desired agent count, before the children are mutated, and the best agent placed back into the agent pool to prevent any fitness degradation as the generations went on.
In this version, first the agents have the score modification from above applied to their final score, before being sorted by their scores, the bottom half of the agents are removed.
The remaining, top scoring half of the population, are added to the new population.
The entire top scoring have are also duplicated, but mutated, as if they had children.
The single best scoring agent is also saved as json, storing its weights, what round it was a part of, its final score, and every input and output for each iteration, this is used in the second part of this project
I don't keep track of it in this project, but this is somewhat like the development of several, separate, species, that may end up becoming the majority of the population.
The actual mutation is relatively simple, a random number between 0-1 is chosen, and if it is less than the given mutation rate, the current weight has a value of -2 - 2 added to it, very simple, but effective.
I originally planned to force at least one mutation per generation, but this usual resulted in performance degradation, as the mutations were usually harmful.
A note on the species
Just as a little experiment, each agent is colored using the first weight value as the hue for the color, allowing some tracking of species, as the children of one kind of agent will probably have a similar color to its parent, one day I do plan on making one full definitive python genetic algorithm project, that follows species, and their lineages, all the way to the beginning of the simulation.
If I end up ever doing that, it would probably be the final evolution of my genetic algorithm projects, as that would meet every single goal I've had about these projects.
A little addition
As I was working on this, I thought that there needs to be some way to visualize the actual network behind the decisions of the agent, as with previous projects, you were able to see each move one by one, I decided to create a tool that lets you visualize the weights, and the inputs and outputs, and the general network easily.

This tool shows the network, and weights, of the agent, the weights are shown as the lines the thicker they are the larger the weight is, if they are green they are positive and red if they are negative.
It shows the inputs and outputs as well, on both sides of the network, and the values of the hidden neurons.
In the bottom right corner, it shows exactly what the agent is seeing, with the white dot being the agent, and the squares representing the nearby tiles.
- Black tiles are walls
- Grey tiles, with just an outline, are empty spots
- Red tiles are apples
- And white tiles are apples that have already been collected
This tool is also interactive, allowing you to use the arrow keys to change the step in the simulation, with the values updating as well, you can also hit the space bar to stop/start playback which automatically goes further through the simulation, step by step, or hit r to return back to the first step.
Clicking on, and selecting a neuron, lets you see all of the actual weights of the connections, that are connected to it.
This can be used by changing the file name variable in view.py and then running it.
A quick analysis
The image from the network visualizer above is from the maze game, showing that the score, and iteration inputs aren't used very much.
I assume this is because the network was able to "learn" the maze somewhat, and use the vision inputs more, to try and guess where in the maze it is.
The up and down values have several negative weights connected to them, I assume this may be because there is more sideways movement vs up and down movement in the maze, so a large weight is placed on the up and down inputs to help stop the agent from running into them.
This image is from the "random" game, where the apples are placed randomly within the area, and the agents need to find as many as possible, meaning that the vision inputs are less useful, as they will only show apples when the agent is directly next to them.

It has more green, positive, weights connected to the movement neurons, and the iteration neuron especially, I assume that this is due to the agent needing to act very quickly if it ever sees an apple, and collect it fast, before it moves away and loses it.
The focus on iterations may allow the agent to have some kind of "search pattern" allowing it to find apples somewhat easier.
This agent also has more positive connections to the bias input, this may just be to ensure some base stimulus, as it will not be getting much from the vision inputs, so it has the ability to move at all, this may also help with the search pattern theory.
Conclusion
This was an interesting project, evolving as it went on! I had no plan when I began to build a full visualizer tool, but I ended up doing so.
In the future I would like to build another genetic algorithm project, that includes species, and the ability to track the lineage of each agent.
Possibly something similar to CaryKH's Evolution Simulator, with a table showing each of the agents.
I also plan on making the next project more interactive.
I also wish to either build a new tool, or modify the viewer, to allow little scenarios to be built, allowing you to not just see the network and its inputs/outputs, but also to modify the inputs, and see how the decisions change based on that.
Generally I wish to make things more interactive in the future, and this was a worthwhile exploration into the technical changes, namely the neural networks and reproduction changes, that would make sense to be added to a definitive version of these projects. To help with the interactivity, I may build the next one in Godot, or in the web languages, to allow for more people to use it.
Thank you for reading, see you next month!
btw happy pride month! and i added a guestbook to my website as well!