Automating the network: what matters before the tools

When a team decides to automate its network, the first question is almost always “which tool should we pick?”. Ansible or Terraform, Nautobot or NetBox, this orchestrator or that one. The question is fair, but it comes too early, and this article explains why. It first shows that operational difficulties come from how the network is operated rather than from the tools, then what an automated network concretely looks like, why so many projects disappoint even though they have good tools, which questions need settling for the choice of tool to become simple, and finally why that choice is about functions rather than tools.
The problem is how the network is operated, not the tools
As long as a network counts a few dozen devices, manual operations seem to be enough. They often are, thanks to two or three people who know everything by heart. Then the infrastructure becomes more distributed, more dynamic, more entangled with cloud and security, and the same symptoms appear everywhere: configurations that drift from one site to the next without anyone knowing why, a simple change that takes days because it touches fifty devices, documentation that describes the network as it was two years ago, critical knowledge that lives in two people’s heads.
These symptoms feed one another. A heterogeneous, poorly documented network makes every change risky. Risk breeds caution, caution blocks the very investment that would reduce the fragility, and the fragility settles in. You end up running the network like an old house: you touch nothing, for fear that everything will come down.
None of these symptoms comes from a missing tool, and none of them goes away because you buy one. They come from the way the network is operated: by hand, from memory, with no shared reference. That way of operating is what automation changes. Before talking about tools, then, we need to describe what it becomes.
What an automated network looks like
An automated network is not a network with scripts running on it. It is a network operated according to four principles, each of which answers one of the symptoms above.
- The intended state is described in one place. The network (sites, devices, addressing, services) is described in a structured way in a Source of Intent, sometimes called a source of truth. That description is the reference, and the network conforms to it, not the other way round. This is the answer to outdated documentation: the description is no longer a document filed next to the network, it is what drives it.
- Every change is treated as code. It is written down, versioned, reviewed by a peer, tested, then applied. You know who changed what, when and why, and you can go back. This is the answer to knowledge living in two heads: it lives in the repository, not in people. Software development has worked this way for twenty years, and that way of working has proven itself.
- A change gives the same result everywhere. On one device or on five hundred, the same operation produces the same state, and rolling back is a planned move, not a rescue operation. This is the answer to the simple change that takes days.
- The gap between intended and actual is measured continuously. The actual state of the network is compared with the intended state, gaps surface before they become incidents, and they trigger an action rather than a report. This is the answer to configurations drifting without anyone knowing.
One could ask what this drift detection is for in an intent-driven approach: since the configuration is regenerated from the intended state, would the next deployment not overwrite the deviations anyway? It is an excellent objection, and the answer comes in two parts. First, trust rests on transparency: automation must announce what it is about to do and account for what it actually changed. Silently overwriting a manual change, without even having flagged the deviation, would contradict that principle. Second, not everything gets pushed again: for speed, some platforms only regenerate and deploy the configurations affected by a change in the Source of Intent. A deviation on a device that nothing touches can then survive for a long time, if nobody measures it.
None of these principles names a tool. A project can honour them with modest tools, and betray them with the most expensive platform on the market. That is precisely what happens in most projects that disappoint: the tool is in place, the principles are not.
Why projects disappoint despite good tools
Three typical situations, composed from what is commonly observed in the field, illustrate that gap. In each of them, the tool is the right one and one of the principles is missing.
The network everyone thought they knew. A multi-site company is convinced its configurations are homogeneous, since they were all deployed from the same template. The first automated audit, read-only, shows that every site has drifted in its own way over years of emergency interventions. Nobody was wrong, nobody had the whole picture. The audit tool invented nothing: it made visible a network that nobody really knew, and on which everyone was about to build.
$ netops audit edge-lyon-01 --diff
--- intended
+++ running
- ntp server 10.0.10.1
+ ntp server 192.168.1.50
- snmp-server community netops-ro RO
+ snmp-server community public RO
interface GigabitEthernet0/1
- description uplink-core-01
+ description TEMP-FIX
✗ 3 differences with the Source of Intent
The script nobody dares to run again. Three years ago, an engineer wrote a script that provisions new sites. It works, but it has no tests and no documentation, it assumes a precise starting state, and its author is the only person who knows which order to run its steps in. The day that person changes team, the script turns from an asset into a risk. The tool was the right one; the automation had simply never been treated as code.
The tool bought before the data. A source of truth has been deployed, the tool is in place, and it is empty. Or worse, it was filled once and never updated again, because nobody owned the data and no process fed it. The tool was the right one here too; nobody had decided who would own the intended state or how it would be maintained.
In all three cases, neither the tool nor the teams’ skills are at fault. What is missing sits upstream: knowing the network before tooling it, treating automation as code, deciding who owns the data. Those are questions, not products, and they get settled before the choice of tool.
The questions to settle before choosing a tool
Here are those questions, in the order they come up. Answering them requires nothing but time spent understanding and designing, and that is precisely the time rushed projects skip to go straight to the tool.
- What does the network look like today, and what should it look like? The starting point is an up-to-date inventory and the actual configurations, not the documentation. But the aim is not to automate the network as it stands, because automation copes poorly with exceptions. Automating a use case has a near-fixed development and maintenance cost: the work is the same for one device or for hundreds. The effort therefore only pays off when it applies to a broad scope. Yet every exception, every divergent pattern, adds another corner case to handle. As they pile up, the automation logic loses readability and robustness, the bill keeps growing, and the return on investment collapses. Automation is therefore an opportunity to simplify the network, by bringing exceptions back to the standard before encoding them. In an existing network you do not start from scratch: you pick a small, well-bounded first scope, begin with the most homogeneous segments, and keep “automated from day one” for new infrastructure. A small scope limits the impact of any mistake, gives the team time to learn, and funds the expansion with its first results.
- Where does the intended state live, and who owns it? The most common trap is letting each tool carry its own description of the network: variables in a playbook, a spreadsheet for addressing, a ticket for VLANs. None of those descriptions is the reference, and they end up contradicting one another. The intended state has to live in one place, in a model that describes the network and how it relates to IPAM, cloud, security and ITSM. That model needs an owner and a process that feeds it, because data quality is almost always underestimated at the start, and a source of truth with no owner empties itself within six months. That model is what will later decide the tool: you look for the one that serves it, not the other way round.
- How will the automation be built, and who will be able to use it? Automation is built like software because it carries the same risks: at scale, a mistake does not hit one device but hundreds. That imposes a few practices, each of which prevents a specific accident. The code lives in a versioned repository, so that you know what changed and can go back. Every modification is reviewed by a peer, so that a logic error is seen by two people rather than one. Automated tests check the data and the workflows before production, because whatever was not tested beforehand will be tested by production. A simulation mode shows what a run would change without applying anything, so that the operator approves with full knowledge. Documentation lets people other than the authors use it, without which the automation is just a new point of fragility. What remains is to frame its use with simple governance: which operations are eligible for automation and which keep a human approval, who can trigger what on which scope, and what record is kept of every run.
- Who carries it, and how will you know it has succeeded? Successful automation rests on three pillars, people, process and technology, and technology is the last one. People first, because network engineers need to understand what is being built, be trained on it, and above all be its authors rather than its audience: automation that is imposed is automation that gets worked around. Process next, because every automation must answer an identified business need and be measurable. You need to decide before starting what you will measure, for instance the time to production for a change, the change success rate, configuration drift or the time to resolve incidents, rather than finding out afterwards.
- What do you buy, and what do you build? This is the only question where the tool finally appears, and by this point it is almost easy. The rule is to buy when you can and build when you must. The generic pieces, orchestration, pipelines, data storage, observability, exist off the shelf and need not be reinvented. The layer specific to your network, on the other hand, data model, templates, workflows, validation rules, cannot be bought: that is what has to be built. In both cases, favouring standards, APIs and open-source tools keeps any platform from locking you in and lets you replace each building block without starting over. Tools change; the data model and the practices stay.
Choose functions, not tools
Once these questions are settled, the reflex comes back: “so, Nautobot or NetBox?”. It deserves one more pause, because the right question is still not which tool but which functions. An automation chain is not a tool. It is a set of functions handing data to one another, and each product on the market covers one or several of them, rarely all.
The Network Automation Forum has formalised that reading in its NAF Framework, a vendor-neutral reference model that breaks network automation down into six functional blocks:
- Intent: describe and store the desired state of the network. That is the Source of Intent of this article.
- Collector: retrieve the actual state from the devices.
- Observability: keep that actual state, process it and surface whatever deviates from the intent.
- Orchestrator: coordinate the execution of tasks, on an event, a schedule or a request.
- Executor: apply the changes to the infrastructure.
- Presentation: give users an interface to interact with the whole.
The loop drawn earlier in this article is a condensed version of this model. The framework does not say which product goes behind each block. It says what each block must be able to do, and it states that one block may be served by several components, or several blocks by a single one. That is exactly what changes the conversation. Instead of comparing tools with one another, you list the functions you need, note those you already have, often more than you think, and then look, function by function, for what fills each one best. A tool that covers three blocks is not necessarily better than one that covers a single block very well, nor the reverse: it depends on the functions you are missing.
This vocabulary has another merit: it shows the gaps. Many projects have an Executor, a set of playbooks, and nothing else: no Intent, no Collector, no Observability. They have a tool; they lack functions. The three situations above read the same way. The network everyone thought they knew had no Collector and no Observability. The script nobody dares to run again was an Executor with no Orchestrator to sequence its steps and no Presentation to let anyone but its author use it. The tool bought before the data was an Intent block with no process to feed it.
The right tool is the one that serves a model, a team and a process already in place, and that fills a function named before it. Settle those questions first; the choice of tool then becomes almost obvious.
Recommended reading: Designing Network Automation at Scale by Christian Adell, Modern Network Observability by David Flores, Christian Adell and Josh VanDeraa, and the NAF Framework by the Network Automation Forum.