An angry admin shares the CrowdStrike outage experience

lemme in@lemm.ee · 1 year ago

An angry admin shares the CrowdStrike outage experience

Justin@lemmy.jlh.name · 1 year ago

Why the fuck does an antivirus need a kernel driver

SapphironZA@sh.itjust.works · 1 year ago

Because the windows OS is inherently insecure with lots of permission elevation opportunities.

Blaster M@lemmy.world · 1 year ago

Pretending linux privelege escalation doesn’t exist… to fight something that gets root you have to be able to fight at the root level, or the root access malware can simply nuke the av from userland.

Justin@lemmy.jlh.name · edit-2 1 year ago

Or you could just use kernel namespaces, SELinux, Systemd sandboxing, etc. There is zero need to run in ring 0 for security reasons.

Also, privilege escalation is a lot rarer on Linux than it is on Windows.

catloaf@lemm.ee · 1 year ago

Because that’s where filesystem access lives? AV wouldn’t do very much good if it could only run from userspace.

db0@lemmy.dbzer0.com · 1 year ago

Pity the administrators who dutifully kept a list of those keys on a secure server share, only to find that the server is also now showing a screen of baleful blue.

Lol, can you imagine? It empathetically hurts me even thinking of this situation. Enter that brave hero who kept the fileshare decryption key in a local keepass :D

kescusay@lemmy.world · 1 year ago

Seems like an argument for a heterogeneous environment, perhaps a solid and secure Linux server to host important keys like that.

pearsaltchocolatebar@discuss.online · 1 year ago

Linux can shit the bed too. You need to maintain a physical copy.

Voroxpete@sh.itjust.works · 1 year ago

Their point is not that linux can’t fail, it’s that a mix of windows and linux is better than just one. That’s what “heterogeneous environment” means.

You should think of your network environment like an ecosystem; monocultures are vulnerable to systemic failure. Diverse ecosystems are more resilient.

noobface@lemmy.world · 1 year ago

Hey Ralph can you get that post-it from the bottom of your keyboard?

gnutrino@programming.dev · 1 year ago

Sure but the chances of your Windows and Linux machines shitting the bed at the same time is less than if everything is running Windows. It’s exactly the same reason you keep a physical copy (which after all can break/burn down etc.) - more baskets to spread your eggs across.

pearsaltchocolatebar@discuss.online · 1 year ago

Very few businesses are going to spend the money running redundant infrastructure on two different operating systems. Most of them won’t even spend the money on a proper DR plan.

Revan343@lemmy.ca · 1 year ago

Then they get to suffer the consequences when shit like this happens

stringere@sh.itjust.works · 1 year ago

Then they get to suffer the consequences when shit like this happens

Oh, they are.

StaySquared@lemmy.world · 1 year ago

CS did take down Linux a few years back… I forget the exact details.

Avatar_of_Self@lemmy.world · 1 year ago

Yes, but has it taken both OS’ out at the same time? It hasn’t but it could happen, however, the chances are even less. There’s obvious risk mitigation in mixing vendors in infrastructure for both hardware and software in the enterprise.

If some critical services were lost in your enterprise last time until RH updated their kernel then you could have benefitted from running that service from Windows as well. Now the reverse is true. You could have another DC via Samba on Linux in your forest if you wanted to, in order to have an AD still for example. Same goes for file share servers, intermediary certificate servers (hopefully your Root CA is not always on the network) and pretty much most critical services.

Most enterprises run a lot of services off of a hypervisor and have overhead to scale (or they are already in a sinking ship), so you can just spin up VMs to do that. It isn’t as if it is unreasonably labor intensive compared to other similar risk mitigation implementations. Any sane CCB (obviously there are edge cases but we are talking in general here) will even let you get away without a vendor support contract for those, since they are just for emergency redundancy and not anywhere near critical unless the critical services have already shit the bed.

Amanda@aggregatet.org · 1 year ago

Sounds like we may have an easier conclusion to draw here

sugar_in_your_tea@sh.itjust.works · edit-2 1 year ago

That’s why the 3-2-1 rule exists:

3 copies of everything on
2 different forms of media with
1 copy off site

For something like keys, that means:

secure server share
server share backup at a different site
physical copy (either USB, printed in a safe, etc)

Any IT pro should be aware of this “rule.” Oh, and periodically test restoring from a backup to make sure the backup actually works.

IphtashuFitz@lemmy.world · 1 year ago

We have a cron job that once a quarter files a ticket with whoever is on-call that week to test all our documented emergency access procedures to ensure they’re all working, accessible, up-to-date etc.

Empricorn@feddit.nl · 1 year ago

Are you hiring!?

MrNesser@lemmy.world · 1 year ago

Lemmy appears to be weathering the storm quite well…

…probably runs on linux

KairuByte@lemmy.dbzer0.com · 1 year ago

I doubt many Lemmy servers are running enterprise level antivirus.

RBG@discuss.tchncs.de · 1 year ago

It runs on hundreds of servers. If any of them ran windows they might be out but unless you got an account on them you’d be fine with the rest. That’s the whole point of federation.

Grandwolf319@sh.itjust.works · 1 year ago

I’m so proud of this community!

cygnus@lemmy.ca · 1 year ago

The overwhelming majority of webservers run Linux (it’s not even close, like high 90 percent range)

Bilb!@lem.monster · 1 year ago

I wonder if any Lemmy servers run on Windows without WSL. I can’t think of any hard dependencies on Linux, so it should be possible.

Boozilla@lemmy.world · edit-2 1 year ago

If you have EC2 instances running Windows on AWS, here is a trick that works in many (not all) cases. It has recovered a few instances for us:

Shut down the affected instance.
Detach the boot volume.
Move the boot volume (attach) to a working instance in the same region (us-east-1a or whatever).
Remove the file(s) recommended by Crowdstrike:
Navigate to the C:\Windows\System32\drivers\CrowdStrike directory
Locate the file(s) matching “C-00000291*.sys”, and delete them (unless they have already been fixed by Crowdstrike).
Detach and move the volume back over to original instance (attach)
Boot original instance

Alternatively, you can restore from a snapshot prior to when the bad update went out from Crowdstrike. But that is not always ideal.

Defaced@lemmy.world · 1 year ago

A word of caution, I’ve done this over a dozen times today and I did have one server where the bootloader was wiped after I attached it to another EC2. Always make a snapshot before doing the work just in case.

Boozilla@lemmy.world · 1 year ago

Good advice!

gravitas_deficiency@sh.itjust.works · 1 year ago

Lmao this is incredible

Another Redditor posted: "They sent us a patch but it required we boot into safe mode.

"We can’t boot into safe mode because our BitLocker keys are stored inside of a service that we can’t login to because our AD is down.

“Most of our comms are down, most execs’ laptops are in infinite bsod boot loops, engineers can’t get access to credentials to servers.”

N.B.: Reddit link is from the source

I hope a lot of c-suites get fired for this. But I’m pretty sure they won’t be.

MagicShel@programming.dev · 1 year ago

C-suites fired? That’s the funniest thing I’ve heard yet today. They aren’t getting fired - they are their own ass-coverage. How can they be to blame when all these other companies were hit as well?

I guess this is a good week for me to still be laid off.

Codex@lemmy.world · 1 year ago

Our administrator is understandably a little bitter about the whole experience as it has unfolded, saying, "We were forced to switch from the perfectly good ESET solution which we have used for years by our central IT team last year.

Sounds like a lot of architects and admins are going to get thrown under the bus for this one.

“Yes, we ordered you to cut costs in impossible ways, but we never told you specifically to centralize everything with a third party, that was just the only financially acceptable solution that we would approve. This is still your fault, so we’re firing the entire IT department and replacing them with an AI managed by a company in Sri Lanka.”

Evotech@lemmy.world · 1 year ago

Stupid argument though, honestly just chance that crowdstrike was the vendor to shit the bed. Might aswell have been set. You should still have procedures for this

SkybreakerEngineer@lemmy.world · 1 year ago

Fired? I hope they get class-actioned out of existence as a warning to anyone who skimps on QA

slacktoid@lemmy.ml · 1 year ago

Sounds like the best time to unionize

gari_9812@lemmy.world · 1 year ago

Any time is a good time to unionize

slacktoid@lemmy.ml · 1 year ago

Agreed, just here they have then by the metaphorical balls.

ɔiƚoxɘup@infosec.pub · 1 year ago

I’m in. This world desperately needs an information workers union. Someone to cover those poor fuckers in the help desk and desktop support as well as the engineers and architects that keep all of this shit running.

Those of us that aren’t underpaid are treated poorly. Today is what it looks like if everybody strikes at once.

slacktoid@lemmy.ml · 1 year ago

This dude here coming in hot with a name, Information Workers Union (IWU). Love it

Soo are you gonna create the community or am I?

ɔiƚoxɘup@infosec.pub · 1 year ago

I’m too chicken.

JasonDJ@lemmy.zip · 1 year ago

Nah.

Bureau of Information & Technology Servicers.

uis@lemm.ee · 1 year ago

BITechS

JasonDJ@lemmy.zip · 1 year ago

Or just BITS.

EnderMB@lemmy.world · 1 year ago

To preface, I want to see a tech workers union so, so bad.

With that said, I genuinely don’t believe that most tech workers would unionize. So many of them are brainwashed into thinking that a union would dictate all salaries, would force hiring to be domestic-only, or would ensure jobs for life for incompetent people. Anyone that knows what a union does in 2024 knows that none of that has to be true. A tech union only needs to be a flat fee every month, guaranteed access to a lawyer with experience in your cases/employer, and the opportunity to strike when a company oversteps. It’s only beneficial.

Even if you could get hundreds of thousands of signatories, the recent layoffs have shown that tech companies at the highest level would gladly fire a sizable number of employees if it meant stamping out a union. As someone that has conducted interviews in big tech, the sheer numbers at peak of people that had applied for some roles was higher than the number of active employees in the whole company. In theory, Google could terminate everyone and replace them with brand-new workers in a few months. It would be a fucking mess, but it (in theory) shows that if a Google or Apple decided that it wanted no part of unions they could just dig into their fungible talent pool, fire a ton of people, promote people that stayed, and fill roles with foreign or under-trained talent.

slacktoid@lemmy.ml · 1 year ago

I feel you with this. They do not see themselves as workers. Thank you for the preface.

EnderMB@lemmy.world · 1 year ago

Agreed, sadly to many there is still the view of tech being a meritocracy, and that they’re in FAANG because of their hard work over everything else, so fuck everyone else. Naturally, many change their tune once their employer actions regressive policies, but it’s surprising how many people just have zero understanding of what a union does. They see cop shows or The Wire and assume it’ll be like the unions there…

rozodru@lemmy.ca · edit-2 1 year ago

deleted by creator

7rokhym@lemmy.ca · 1 year ago

Just a thought from experience: Be wary of any critical products and/or taking a job from a company run by an accountant. CrowdStrike CEO… accountant!

Accounting firms are an obvious exception.

Buffalox@lemmy.world · edit-2 1 year ago

At leas no mission critical services were hit, because nobody would run mission critical services in Windows, right?
…
RIGHT??

OpenStars@discuss.online · 1 year ago

scottywh@lemmy.world · 1 year ago

If it only impacts a percentage of your machines then there was a problem in the deployment strategy or the solution wasn’t worthwhile to begin with.

Phoenixz@lemmy.ca · 1 year ago

… So your point was that it would have been better if everything went down?

There are plentiful reasons why deployments are done in parts, and I’m guessing that after today strategies will change to apply updates in groups to avoid everything going down.

Also, dear God, stop using windows as a server, or even a client for that matter. If you’re paying actual money to get this shit then the results are on you.

scottywh@lemmy.world · 1 year ago

Also, 😂

scottywh@lemmy.world · 1 year ago

No.

My main point was that crowdstrike has always been lazy man’s garbage.

Hurculina Drubman@lemm.ee · 1 year ago

I got super lucky. got paid for my car just before the dealership systems went down, got my return flight 2 days before this shit started.

Cossty@lemmy.world · 1 year ago

I didnt know so many servers still run windows.

stoly@lemmy.world · 1 year ago

I can’t imagine how much work it would be to migrate all your services onto Linux. The problem was people adopting windows in the first place.

douglasg14b@lemmy.world · 1 year ago

I love the Linux bros coming out of the woodwork on this one when this could have very well have been Linux on the receiving end of this shit show. Given that it’s a kernal level software issue, and not necessarily an OS one.

It’s largely infeasible to use Linux for many, most, of these endpoints. But facts are hard.

flop_leash_973@lemmy.world · edit-2 1 year ago

They are just butt hurt that this whole thing really shines a light on how inaccurate the line of “the world runs on Linux” truly is.

The world runs on a lot of different things for different reasons and that does not fit nicely into their Richard Stallman like world view.

pathief@lemmy.world · 1 year ago

Just to clarify: the world runs in linux servers. The market share for the non-server market is abysmal.

jabjoe@feddit.uk · 1 year ago

Except lots of IoT things, router, etc. Also Cromebooks and Steamdecks. And us GNU/Linux people. Android is Linux, just not GNU/Linux. Really isn’t just servers.

pathief@lemmy.world · edit-2 1 year ago

I don’t know too much about IoT but I wouldn’t say linux runs the world in any of the other markets you mentioned.

I would say while technically Android uses a modified linux kernel, you can’t put it under the same umbrella.

Either way I don’t want to get too much into these technicalities. I was simply trying to say that Linux is king on servers, not really on the market where all this crazyness happened.

jabjoe@feddit.uk · 1 year ago

Android makes RMS’s GNU/Linux language make sense. It is a Linux, but not a GNU/Linux.

Google’s attempt to fork Linux failed and now is mainlined so they can maintain as small a set of patches as they can. Once binder was merged, there is no fork anymore.

https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/tree/drivers/android?h=v6.10

Android is basically a build config now.

TVs that aren’t Android are probably GNU/Linux. Smart white goods are often Linux. Linux even get used in cars. Some of it under Automotive Grade Linux, but not all. If some random thing has a user interface, find licenses and you can normally see what FOSS went in.

You can do so much with so little, at no cost of licencing or access. Why wouldn’t you?

You use things like Yocto and Buildroot to build a image that has nothing but what you need, how you need it.

theonyltruemupf@feddit.de · edit-2 1 year ago

🐧: what is my purpose?
🧑‍💼: you run web servers.
🐧: oh my god

EdgelDil@lemmy.world · 1 year ago

I appreciate that rick and morty ref

jabjoe@feddit.uk · 1 year ago

The is no single Linux. It’s not a monoculture like that. There are many distros with different build options, different configurations and different components.

Also culture is different. Very few Linux admins would be happy putting in a closed blob kernel driver for anything. In Windows world that’s the norm, but not Linux.

What’s just happened to Windows world would be harder in Linux world. At worse, one distros rolls out a killer update. Some distros would just reboot to the previous kernel.

save_the_humans@leminal.space · edit-2 1 year ago

Hey man, let us have this one. Any immutable/atomic distribution could have either prevented this or easily rolled back the update. Not to mention a Linux offering by something like Red Hat, for example, wouldnt recommend installing closed source third party kernel modules for exactly this reason. Not sure about the feasibility of these endpoints, but the way things are generally done on, and the philosophy of, Linux could very well have avoided this catastrophe.

stephen01king@lemmy.zip · 1 year ago

Can you explain what is immutable/atomic distribution and how it can prevent this?

Captain Aggravated@sh.itjust.works · 1 year ago

An immutable distribution is one that treats the system files as read-only. Applications are handled separately, and updates to the system are done in an image-based way, rather than changing a few updated files, basically the OS gets replaced with an updated version. It prevents users or malicious outsiders from just changing system files. Fedora Silverblue and SteamOS as found on Valve’s Steam Deck are examples of immutable distros.

Now, with soemthing like Crowdstrike that operates in kernel space…I’m too far outside my wheelhouse to grasp how that would work on an immutable system. How it would be implemented.

kent_eh@lemmy.ca · 1 year ago

My former employer had a bunch of windows servers providing remote desktops for us to access some proprietary (and often legacy) mission critical software.

Part of the security policy was that any machines in the possession of end users were assumed to be untrustworthy, so they kept the applications locked down on the servers.

TonyOstrich@lemmy.world · 1 year ago

I kinda wish my employer would do something like this for our current applications. Right before I started working there they switched from giving engineers desktops to laptops (work station laptops but still). There are some advantages to having a laptop like being able to work from home or use it in a meeting, but I would much prefer the extra power from a desktop. In mind the best of both worlds would be to have a relatively cheap laptop that basically acts as a thin client so that I can RDP into a dedicated server or workstation for my engineering applications. But what do I know ¯_(ツ)_/¯

bruhduh@lemmy.world · 1 year ago

Proxmox with windows containers is used widely

CaptPretentious@lemmy.world · 1 year ago

I’m the corporate world, very much Windows gets used. I know Lemmy likes a circle jerk around Linux. But in the corporate world you find various OS’s for both desktop and servers. I had to support several different OS’s and developed only for two. They all suck in different ways there are no clear winners.

Dark Arc@social.packetloss.gg · 1 year ago

It’s not just a cycley jerk in this case. Worried is dominant for desktop usage but Linux has like 90% of the server market and is used for basically all new server projects.

Paying for Windows licensing when it doesn’t benefit you, it’s silly, and that’s been realized for years.

EnderMB@lemmy.world · 1 year ago

Web servers, sure, but even companies that manage infrastructure like Google and Amazon have a LOT of Windows servers kicking around for shit like AD, Outlook, Federation, Office/Teams, etc.

Alborlin@lemmy.world · 1 year ago

Thank för addressing Lemmy circlejerk för Linux . They really take it far

terminhell@lemmy.dbzer0.com · 1 year ago

On prem AD. At least for my MSP’s clients. Have been pushing hard last few years to migrate to azure.

Hotzilla@sopuli.xyz · 1 year ago

Issue is not just on servers, but endpoints also. Servers are something that you can relatively easily fix, because they are either virtualized or physically in same location.

But endpoints you might have thousand physical locations, and IT need to visit all of them (POS, info/commercial displays, IoT sensors etc.).

Miaou@jlai.lu · 1 year ago

Parent comment applies even more so to such endpoints imo

Knock_Knock_Lemmy_In@lemmy.world · 1 year ago

Do IoT sensors really run windows? If not, how are they affected?

Disaster@sh.itjust.works · 1 year ago

80% of our machines were hit. We were working through 9pm on Friday night running around putting in bitlocker keys and running the fix. Our organization made it worse by hiding the bitlocker keys from local administrators.

Also gotta say… way the boot sequence works, combined with the nonsense with raid/nvme drivers on some machines really made it painful.

TheObviousSolution@lemm.ee · 1 year ago

It might be CrowdStrike’s fault, but maybe this will motivate companies to adopt better workflows and adopt actual preproduction deployment to test these sort of updates before they go live in the rest of the systems.

EnderMB@lemmy.world · edit-2 1 year ago

I know people at big tech companies that work on client engineering, where this downtime has huge implications. Naturally, they’ve called a sev1, but instead of dedicating resources to fixing these issues the teams are basically bullied into working insane hours to manually patch while clients scream at them. One dude worked 36 hours straight because his manager outright told him “you can sleep when this is fixed”, as if he’s responsible for CloudStrike…

Companies won’t learn. It’s always a calculated risk, and much of the fallout of that risk lies with the workers.

MrAlternateTape@lemm.ee · 1 year ago

That comment about sleep…that’s about where I tell them to go fuck themselves. I’ll find a new job, I’m not going to put up with bullshit like that.

uis@lemm.ee · 1 year ago

Sounds so illegal, that it makes labour authoririty happy

EnderMB@lemmy.world · 1 year ago

Is it illegal? I’m not American so I have no idea if there are laws in your country against on-call maximum hours.

uis@lemm.ee · 1 year ago

It’s not about oncall, they are literally in the office
See 1
Not sure about America, but it is very illegal in Russia.

Disaster@sh.itjust.works · 1 year ago

That dude should not have put up with that.

Randelung@lemmy.world · 1 year ago

Oh sweet summer child.

cheetah_cheetos@lemmy.world · 1 year ago

Might be hard to do. Crowdstrike release several updates per day to the channel files to match changes in adversarial behaviour. In this case, BCP and backup are what need to be done.

Fedizen@lemmy.world · 1 year ago

lmao