DISCLAIMER: This is an opinion piece first and technical guidance second. The opinions expressed here are all mine, and they are only partially based on hard data. It is therefore perfectly possible that the true reasons why the ESAE architecture didn’t get the traction it has always deserved are entirely different from what I am about to lay out here. But I have been doing this long enough to be confident that I am probably not too far off the mark. If you share the opinions, the technical guidance expressed here is for you.
What is ESAE anyway?
ESAE, or Enhanced Security Administration Environment, is a combination of three design components:
- Red Forest (initially, it was simply another, hardened forest with a one-way trust, since ESAE was created pre-2016)
- Administrative Tiering (sporting a fixed three-tier model based largely on URA, since ESAE was created pre-2012)
- Privileged Access Workstations (they necessarily follow from the other two items, I will explain later)
The adoption, both among customers who got this architecture recomended directly by Microsoft and those who were guided by third-party – or even inhouse – experts, fell way short of the expectations. Here’s where I would love to have hard data in form of real feedback from real admin and security teams, delivered in a timely manner after an EASE project was concluded, abandoned or rejected outright.
What I believe some/a few/many/most consultants, both from Microsoft and from independent parties, have failed to convey to customers about why the ESAE model really can help them get more secure, is this.
The principles of ESAE are not some universal tenets of secure IT environment design. They address very concrete product-specific weaknesses in Active Directory and Windows. An organization needs ESAE not because the world (of IT) needs ESAE, but because the organization had decided to base its IT infrastructure on Windows and its identity delivery infrastructure on AD.
Conversely, there are no other first-party means to address those challenges, and a possible third-party intervention must go so deep into the bowels of Windows that it will most certainly introduce disruption, on top of monetary cost, learning curve and possibly broken backwards compatibility.
There have been lots of misconceptions going on, about whether parts of ESAE can be left out without making the other parts utterly useless, about implementation techniques for the components involved – all the way up until whether ESAE is actually „retired“, „deprecated“ or „still the way to go“. I hope to clear some of these up and to show you a way forward, towards a more secure Windows-based IT infrastructure.
The motivation for ESAE
The intrinsic weaknesses that ESAE has set out to mitigate are these:
- In an AD forest, any principal can be made God by assigning privileges and applying policies. As long as the potential attacker (i.e. a normal user who has been social-engineered or a normal endpoint that has been pwned) and the most privileged assets are all in the same forest, elevation to God is technically possible. The obvious solution within the confines of the chosen technology is to NOT have the most privileged assets in the same forest. Hence, Red Forest.
- An interactive (or service, or batch) logon session on a Windows machine leaves credential artifacts on disk and in memory, both during the actual session and after it has been terminated. If the machine in question has been pwned, the attacker may be able to acquire these credentials and reuse that material either to crack the cleartext password, or even directly (pass-the-hash et al.). We cannot prevent a pwned machine from revealing credentials of its daily user to the attacker, so the obvious solution to at least protect privileged credentials from this type of situation is by preventing privileged principals from logging on to machines that are more likely to get pwned. Enter Administrative Tiering.
- An authenticated network connection from a privileged machine (or, more generally, made by a privileged principal) to a pwned machine can reveal credential material to an attacker or enable them to abuse the authenticated network session by relay, replay, or Kerberos delegation. This is why Administrative Tiering must also prevent unprivileged users from logging on to privileged systems – while their credentials are not at risk in a more secure zone, there can still be harm coming from such a session. I wrote about it two years ago in the context of SSH on Windows Server.
- The policy that privileged principals should not be allowed to log on to pwnable endpoints begs the question: where should they be able to log on? Infrastructure servers and RDP-Jumphosts are not the answer to the question, because there still must be an endpoint for the human to sit in front of. And that is how we arrive at the concept of Privileged Access Workstations.
The principles of ESAE are, therefore, just applying logic to the non-changeable constraints of the chosen technology stack.
So why didn’t ESAE conquer the Windows world?
I have always liked to quote the Ten Immutable Laws of Security, but the universal law that is most pertinent here is #2 from the follow-up Ten Immutable Laws of Security Administration:
Security only works if the secure way also happens to be the easy way.
And this is where the initial ESAE blueprint fell flat. All three parts of it did, to a varying degree. The (operational) cost of implementing the ESAE architecture became apparent to the decision-makers very quickly; its security benefits rarely did.
With the Red Forest part, we (as in: those who recommended the architecture to businesses, that includes Microsoft, independent consultants and internal security forward-thinkers alike) simply failed to explain to the Management that we are not introducing a completely new management paradigm here. On the contrary, the beauty of the Red Forest design is that 100% of the technology being used is already intimately familiar to the IT team, while at the same time the attack surface of the most privileged principals that are actually being used is reduced quite drastically. The privileged principals that are not normally being used, i.e. the Break-Glass accounts, remain in their respective Golden Forest, but they are already protected from many attack vector by the virtue of not authenticating to the domain „at peacetime“!
Tiering got caught between IT generalists having to juggle at least four accounts in their day-to-day, with password rotation for every one of them governed by a „Burr Policy“ with its short change cadence and untypeable complexity, and application owners and Service Desk personnel unable to meet their KPIs because of the perceived inability to actually get anything done with this new admin accounts of theirs. But many organizations didn’t even get to experience this aggravation, the tiering initiative forever stuck in the quagmire of the endless design discussions. Ironically, the question whether a certain system is T1 or T2 is usually much harder to answer than drawing the line between T0 and the rest. I will get to this in a second.
Don’t even get me started on the PAWs. The principles of PAW design (clean mainboard + clean keyboard + clean OS + clean network) are solid and good. However, for some reason, IT people who are usually pretty good at projecting external requirements onto internal processes all of a sudden lost all power of interpretation and started applying those principles as the gospel, the only party benefitting from it being the notebook industry. On the other end of the spectrum we, to this day, find „Tier 0-Jumphosts“ with both connectivity and logon rights wide open and AD administrators routinely assigned – and constantly utilizing! – local admin rights. Here, too, the education of the intended users never did the concept justice, resulting in extremes being tried, sometimes even implemented, but mostly not delivering the desired effect.
What can I do today, in my organizaiton?
Take a step back, reevaluate, and do what’s best for the security posture of your IT landscape. And that should always start with securing the most privileged assets – precisely what ESAE had set out to do.
But first and foremost, don’t make the one mistake that will send you back to square one:
Don’t tell them you do ESAE!
If you find yourself tasked with securing a Windows-heavy environment that is not joined to – and sourcing its security from – „the Cloud“, you will probably have no choice than to use some combination of the ESAE components if you want to achieve that goal. Unless you work for an organization that tends to always follow every Microsoft whitepaper to the letter, you’ll be doing everyone a favor if you omit any mention of ESAE from your project definition. „Isolating privileged administration“ can cover both Red Forest and Tier 0 isolation. I’m sure you will find other keywords from your inhouse cyber-speak that will get you through the door. The moment you mention ESAE, prepare for endless discussions about how the whole architecture is „retired“, „overblown“ etc.
Let’s do Red Forest!
I talked about it in my book, in my talks, and wherever I can find someone who will listen: There is no cheaper way to remove your privileged administration from production than a Red Forest. Even before Server 2016, back when ESAE was first formalized, having a one-way trust from every Golden Forest to a common Red Forest already created a good security boundary between privileged identities and production services (i.e. potential ingress points). Of course, we now know that a full compromise of the Golden Forest could allow the attacker to at least get a foothold (at Authenticated Users level) in the Red Forest, but not being able to simply reach the admin accounts already was a significant initial deterrent. And we now also know how to harden this (hint, hint: Authentication Policies).
With Server 2016 and the „Privileged Access Management“ forest optional feature you get the ability to work with Shadow Principals rather than users or groups, which allows you to maintain administrative groups in the Golden Forest that have high privileges there – but are always empty!
Even without a Red Forest, the PAM feature allows you to maintain administrative groups that are MOSTLY empty – by using Just-In-Time Administration (JIT). Within your Red Forest, JIT can even be used in conjunction with Shadow Principals, minimizing exposure for critical permission levels to the absolutely necessary minimum.
I certainly urge you to design your „modern“ Red Forest around these concepts. If you think they lack proper automation, that’s because these features were created to be used with Microsoft Identity Manager 2016 rather than standalone. Unlike Server 2016, MIM is not going out of support in 2026 but has a longer lifecycle policy with extended support expiring – for now – on January 9, 2029. But with a bit of PowerShell (that’s where you come in) these features are perfectly automatable even without MIM, and after the initial setup, privileged administration frameworks are usually not THAT dynamic.
Let’s NOT Do „Tiering“!
You read it right. A full-scale „Tiering initiative“ is the best way to end up not geting any tiering implemented at all. I have seen it happen more times than I care to remember. There are three things that are inherently wrong with the idea of designing and then implementing a complete tiering model for an entire organization:
- The original tiering model treats escalation from regular user to T2 admin to T1 user to T1 admin the same as escalation into T0 from a lower tier. The reality could not be more different. Most of what can be achieved by elevating to T1 admin is also achievable by lateral movement within T2 – at least as far as data access and exfiltration are concerned. In contrast, elevating into T0 means game over.
- The reasoning behind the original design assumes that any user logged on to a normal T2 workstation will behave the same way and make the same mistakes – while the same person, logged on to a Tx PAW with their Tx admin account will in equal measure know what to do. I would like to call both parts of this assumption into question. The persons trusted with T0 activities are hopefully on a different level in terms of cybersecurity awareness than regular users or even T1/T2 admins (i.e. Service Desk or application owners). There are also much fewer of the former than of the latter within the IT organization, and even fewer of them are persons external to the organization where education and control may be difficult.
- The original tiering model recommends, and in some of the documentation requrires, that all three tiers be designed, managed and secured in the same way, because uniformity reduces complexity, and that makes the whole thing more manageable. While uniformity does reduce complexity, the security zones in any IT environment play by very different rules, depending on how they are defined and what their respective impact is if breached. And here we must acknowledge the stark contrast between Tier 0 and everything else. T0 is something that is defined through technology alone. Be honest with yourself, and you can map out your T0 pretty quickly. Everything below that is largely a matter of definition, which is why you will spend way more time separating T1 from T2 than on separating out T0 from everything else – with a very limited positive impact on security posture. Besides, the further away you are from T0, the more „unknown unknowns“ you will encounter.
So the prospects of doing all tiering at once on the drawing board and then pouring the result out into production are rather bleak. But here’s how you will succeed.
Start with Tier 0 isolation
I already mentioned that Tier 0 is defined through technology. What I mean by this is, whether a system or a principal is Tier 0 does not have anything to do with it being classified as Tier 0, or with formal „tiering“ being in place to begin with. This makes determining your effective T0 perimeter a relatively easy task. The good people at SpecterOps have done some great work that will help get you started: https://github.com/SpecterOps/TierZeroTable. Semperis offers ForestDruid free of charge – a great little tool that will also help with determining your true T0. Most probably, you will not like the result of the first exploration, but remember that if you find a T0 object that should not be T0, no amount of discussion and paperwork will change that – only removing the control paths into T0 can do it. So go clean those up.
I do not intend to provide a complete breakdown of how to properly isolate Tier 0 in this blog post. This is coming as a chapter from THAT OTHER BOOK, but the most important design principles here are:
- Apply protection to the objects being protected, as far as technically possible. In terms of technology, it means using AuthN Policies first, and User Rights Assignment policies second – because URA has to be applied to T1/T2 systems, i.e. the very systems we have to assume being compromised.
- Use the very thing that makes an object T0 to determine where to apply protection. That doesn’t work with AuthN Policies, those have to be assigned explicitly, but everything else should be targeted by administrative groups and linked on the domain level rather than purely by OU placement.
If you manage to isolate T0 administration, you have already achieved a significant reduction of your attack surface, and probably not many people outside of the elite group of T0 admins have been inconvenienced by this. After your core team was able to gather enough day-to day experience operating in a tiered infrastructure within your concrete IT environment, you are ready to assist the rest of your IT organization in implementing separation of duties in their respective areas.
Underneath Tier 0, implement zones rather than tiers
On the surface, there is no difference between „zones“ and „tiers“ – both concepts represent groupings of computers and users configured in a way that the user accounts that belong to a particular zone are not allowed to log on interactively to computers outside that same zone. Underneath Tier 0, we are not including the reverse requirement in the overall definition – it may be imposed on certain zones, but not on others. The relationships between zones and tiers is this:
Two zones belong to the same tier if a control path from either zone into the other zone constitutes lateral movement.
Two zones belong to different tiers if a control path from one of the zones to the other constitutes privilege elevation.
This definition, by the way, also allows for multiple zones within Tier 0 – a design I highly recomend for larger and more complex environment where only a smal part of the T0 admin accounts are allowed to log on to actual Domain Controllers and Certtification Authorities classified as T0.
Why should we be rather doing zones in the T1/T2 area and how would we define their boundaries? The answer to the second question is, by application or business process. Imagine an app that relies on Windows authentication while doing its own authorization, and whose client piece also includes the administration interface for that application. This used to be a quite widespread architecture – and in fact, it still is, if the application in question is both used and administered by means of a web frontend. Under these circumstances, the workstations sanctioned for use of the application in question belong in the same security zone as the application’s backend. Of course, zoning off such an application is only going to improve the overall security posture if this application is a possible attack target – because of the data it stores (intellectual property, personal information or commercial data) or because of the business processes it is able to initiate (money transfers or other transactions). And this is the answer to the first question and also one of the reasons why rigidly tiering apart workstations from application servers was never going to be understood by real IT people supporting real businesses.
To add to the motivation to NOT do the rigid T1/T2 separation as per the original ESAE guidance, there is no evidence whatsoever that escalation to Tier 0 from a workstation is any more difficult than escalation from an application server in case appropriate credentials could be harvested from the endpoint in question. As soon as T0 has been properly isolated, the remaining attack surface becomes much more homogenous from the identity and infrastructure security perspective, so that further isolation decisions must be moltivated by possible business impact rather than escalation of privileges.
The initial rigid three-tier guidance seems to suggest a proper separation can never be fully implemented, so that a T1 operator remains a higher escalation risk because people are people and will continue insisting on logging on where they shouldn’t. If you can prove that you can indeed achieve a proper T0 isolation, the importance of an impenetrable border between T1 and T2 sinks dramatically.
Let’s do PAW but not at the missile silo level!
Of course, if your organization does operate a missile silo or two, you have my blessing to follow the strictest security guidance you can find to the letter. But everyone else should try to find balance between security and manageability, see Law #2 quoted above.
Most importantly, look into achieving the Holy Grail of PAW deployment – a virtual PAW running on the Tier 2 notebook. The reason we have all been reluctant to „just put it on there“ is partly that in our collective mind all T2 notebooks are created equal. However, it is absolutely within our sphere of influence as T0 admins to restrict who is allowed to use a particular T2 notebook. Guess what – it will end up being the most security-savvy and responsible people within our IT organization. Namely, persons trusted with a T0 admin account. I am of course making assumptions here, but the assumption that this target audience should be more comfortable with Application Control, BitLocker + PIN and other reasonable hardening techniques than the average user seems solid to me.
The „Clean Mainboard“ tenet you can only satisfy by procuring your notebooks from a trustworthy source, which, in a business setting, should not be a question anyway, even for normal run-of-the-mill workstations.
The „Clean Keyboard“ directive is part physical (no unknown devices plugged into USB) and part digital (no keylogger-type malware injected into the OS). The physical part is, again, something that common sense and workplace hygiene should take care of for everyone, and the digital part is where EDR, Application Control and BitLocker (with PIN!) come into play. Most importantly, though, DO NOT GIVE ADMIN RIGHTS TO ANYONE ON THESE LAPTOPS. LAPS is great – for all the other machines. The laptops carrying the payload of a PAW VM must be 100% managed. And if you end up with a bunch of custom actions in your client management just for these – so be it.
Actually, no. LAPS is not great, it’s just 1000x better than having the same local admin password on the entire client fleet which is the real-world alternative. But using LAPS still means giving admin rights to people who have no business having them, even if it’s „temporary“. Because if criminal energy – or a drive-by download – is involved, there is no such thing as „temporary admin rights“.
Of course, there cannot be local admin rights on the PAW VMs either. But that should not need an explanation. And once you have hardened host AND guest to a degree you can classify it as a „Clean Keyboard“ infrastructure, you have already got „Clean OS“ figured out!
This gets a bit more involved than this short section of a long blog post, but you get the gist – it’s doable. Sami Laiho does a masterclass on this at CQURE, and I encourage you to attend it – you will learn more from a guided session with Sami than you possibly can from reading anything I can put out. Nevertheless, I will be covering this topic in an upcoming chapter of THAT OTHER BOOK, so start wherever it is convenient for you, but start now.
Detailed guidance to be found…
I have started describing these, and other related techniques in greater detail, on my book’s website in the THAT OTHER BOOK section. The first chapter, on Authentication Policies, is already out, the ESAE-related ones are in the queue and will appear in the coming weeks/months. I will put direct links in the sections above once that content is published.
Happy Hardening!