Rendered at 14:44:11 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
stmw 1 days ago [-]
Despite all of the snark here, in my experience Salesforce SRE team is quite competent. The engineering challenges of running a large PaaS - not just with own apps, but with millions of customer-written apps running on it - are quite interesting, and sadly things happen. The status page makes sense to actual customers, it's the particular "pods" where a given service runs.
Anon1096 24 hours ago [-]
Hacker News is much easier to read when you realize that 95% of people have never worked on a "high" (maybe we could say >1B requests per day as a starting point) scale distributed service and think it's trivial to run one with more than 2 nines. You see comments all the time here mentioning that their own desktop at home is achieving more than that which belies deep misunderstanding of how systems are measured. Or that unofficial github status page repeatedly posted here that counts all github services together into one number.
21 hours ago [-]
leptons 23 hours ago [-]
While not a home-run server, the NTP system is a distributed service that receives 100 billion to trillions of requests per day, and it's running pretty smoothly - it's never gone down completely since it started in 1985. It's also very simple. The reason it has so many 9's uptime is because it is simple. Given a low amount of complexity, it's not unreasonable to think that an individual could run a >1B requests per day service.
Salesforce is not simple. It's wildly, overly complex. It's amazing it has any 9's at all and not 8's or 7's. Salesforce offers three 9's, which allows for 43 minutes downtime per month. The current outage is at 8 hours (and counting) so Salesforce is now at 98.9% uptime for the month - there's an "8" in there now. Not good, but considering the complexity of Salesforce, it's still kind of amazing.
pixl97 21 hours ago [-]
>Salesforce is not simple. It's wildly, overly complex.
It turns out business environments are wildly overly complex.
carefree-bob 18 hours ago [-]
It's the scaling nature of enterprise software. If you have a mature B2C app, you have millions of users. What each user wants isn't so important, so it's more of a take it or leave it experience. If you have an enterprise app, one big company can and does push you around to get their features in. And as you grow, you get a few hundred big companies that push you around. The result is this huge mess of features, and now you have to maintain this mess.
I remember Cisco before iOS used to have hundreds of branches for their router, one branch for each major customer that was demanding specific features. It was unmanageable, but that's what you needed to do to win those "enterprise customers".
It also turns out customers aren't very good at articulating their needs and putting them into a cohesive vision of the product. But they sure have specific demands to get stuff in. I'm not blaming the customer, this is just how this world works -- All of the "enterprise software" apps are extremely complex with hidden knobs and weird behavior that was pushed in by a customer twenty years ago all over the place.
jamesfinlayson 13 hours ago [-]
> I remember Cisco before iOS used to have hundreds of branches for their router, one branch for each major customer that was demanding specific features. It was unmanageable
Ouch. I heard of a company in my home town (small B2B service provider) doing something similar - they paid well but I didn't think it was worth it.
lifeisstillgood 21 hours ago [-]
I think it’s like advertising - 50% of my code is wildly over complicated - I just don’t know which 50%
But the GP is essentially correct - there is a 2% of salesforce that could be built run and keep 80% of salesforce users happy. Except that you could not charge enough to be able to advertise on F1 cars and take SVPs out to dinner.
So you could not actually make 80% of them happy - they would ever buy it.
DANmode 19 hours ago [-]
> I just don’t know which 50%
Yes, you largely do - they’re the commits that get rushed to, and through.
This take that showstopping technical debt is unavoidable is very new, and will age like milk.
fragmede 18 hours ago [-]
> showstopping technical debt is unavoidable is very new
No it's not. The push and pull between shipping and paying down technical debt is as old as there's been software to sell. Sales has been selling features that don't exist quite yet ever since they've been talking to customers, and engineering has been pushing back on implementing them yesterday since there's been features to implement. Showstopping technical debt is merely a side effect of who wins that argument in a given org.
DANmode 17 hours ago [-]
Yes.
My point is the technical debt is stopping the show way more often.
Not “this never existed before selling”,
but “we never had the team in place who could do this right in the first place”.
You can quibble about who is responsible, but the fact remains.
ocdtrekkie 23 hours ago [-]
> which belies deep misunderstanding
I think you are missing the point. When I state my Exchange server is more reliable than Exchange Online, I don't think I'm a better engineer. I recognize Microsoft has harder problems to solve than I do. I think building overengineered, oversized SaaS environments is introducing extreme risk. It's an inherent flaw of the current approach.
Smaller is, in fact, better, because it's easier to operate reliably.
jedberg 23 hours ago [-]
Is it? When your internet is out for five days because your ISP takes a few days to get to you, do you acknowledge that you're now at 98.5% availability for the year, far worse than any SaaS email service?
I think people forget that those large environments are there for a reason. To make sure the service stays up in the face of problems outside your own control.
toomuchtodo 23 hours ago [-]
In my entire adult lifetime (mid 40s), my ISP has never been out for five days. Compare to Github, Microsoft, Salesforce, and AWS outages that are always occurring in some fashion. Reddit is down constantly in various ways and still continues to operate as a business, public no less, so I disagree about the need to chase five nines and broadly speaking, large distributed systems that are potentially unnecessary for the use case and target outcome.
Consider yourself lucky that you’ve never been the victim of a fiber cut. But what about if the power to your house goes out? Or what if your server blows the power supply?
My entire point is that you have no redundancy in your system and you also aren’t big enough to have any pull with the vendors who can fix these types of outages so you’re basically at the mercy of your providers with no recourse.
That’s why these systems are built the way they are.
And generally four nines is considered the gold standard these days. I can tell you for sure that both Netflix and Ebay would lose money anytime they drop below four nines because I have at some point been responsible for both. You’re correct that Reddit has a lot more leeway and outage time before they start losing money but not that much leeway.
torginus 20 hours ago [-]
> That’s why these systems are built the way they are.
Built how? Because I can state with confidence that I have cleaned up a ton of failed upgrades/zombie terraform deploys of these serverless kubernetes wonders that followed every best practice under the sun, and these things are not really considered even moderately reliable (as designed by imperfect mortals under real world conditions), meanwhile professionally, people who stand to lose a lot of money should their systems go down generally operate systems whose architectures were designed decades ago, are generally horizontally scaled monoliths, and are extremely conservative in software choice.
Also downtime often is no biggie, as long as it's planned and or don't lose (too much) customer critical data.
Like nodobody cares if your test db cluster goes down for the weekend. We even shut down our db instances to save money.
fc417fc802 14 hours ago [-]
None of that is an argument against small being more reliable. Rather it's an argument that the smaller you are the more important being distributed becomes when managing mundane day to day failures.
Even if a solar flare takes out an entire continent or two I think it's safe to say that the bittorrent network will still be running in some form. Can you be so certain about any given SaaS product?
toomuchtodo 23 hours ago [-]
I've contributed to building out data centers, as well as managed colos for others, primarily in downtown Chicago at Level3 and at 350 E Cermak. I am familiar with architecture required for reliability and diversity, from power and fiber in all the way up the stack to the Kubernetes cluster and software defined networking. If you participate in the capital markets, your data traverses systems I've participated in designing and implementing. There is a time and place for complexity (in this context, large/global distributed systems), but too often, complexity exists where it need not (imho).
"What are you optimizing for?" is always an important question, as is "The Five Whys."
Semaphor 21 hours ago [-]
I mean, if you really need redundancy, isn’t a second instance on a VPS somewhere that you manually switch over to, enough?
Over multiple ISPs, so far internet outages for more then a few minutes is very rare (though the few minutes would make me not want to host something requiring high availability; and a cut cable is really annoying because there simply is no quick fix), power outages even rares, I experienced 3 in 40 years, and the longest was 6 hours.
chias 20 hours ago [-]
Generally, no.
Ignoring for now how you are synchronizing the database and filesystem, and how doing so may well result in your duplicate experiencing the same failure as the original, you can maybe recover from a small class of availability issues that could knock you out of an SLA.
But that assumes you can get online and can fully orchestrate the transition within less than 53 minutes of it starting. Including the time you took to become aware of it. And including the time to diagnose and decide that a switchover would resolve the problem. Including the time it takes for DNS caches to expire and point to the new host. Including the DNS caches which may ignore your TTL. And including all these things again when you switch back.
And assuming, of course, that it doesn't happen again for a whole year.
ocdtrekkie 22 hours ago [-]
I've dealt with a fiber cut, it wasn't nearly that bad. Fiber cuts impacting my SaaS providers were worse because there was nothing I could do about it.
temp_praneshp 23 hours ago [-]
Curious, what AWS outage has affected you for days?
(I hope you'll agree that the middle east outage is a true outlier)
toomuchtodo 23 hours ago [-]
Indeed, the scale Anon1096 refers to wrt distributed systems is anti pattern. It is designed to vacuum up revenue and create enterprise value with scale, not to create resiliency for customers (although resiliency might be a byproduct of a well architected and operated distributed system at scale).
"Simplicity is the ultimate sophistication." -- Da Vinci
a_conservative 22 hours ago [-]
Hidden in this discussion around self-hosting reliability are other options as well.
Depending on your time and appetite for tinkering with all of this, it's not hard to imagine a home setup that fails over to a cheap Hetzner or DO VM. A manual failover at the DNS level isn't overly complex, and could be scripted.
Keeping a database in sync between home and the instance might be simple or more complex depending on needs, but would it really be that hard to have Claude help you setup a replicating Postgres server? If your database (or data files) are 1 gigabyte and don't update that often... maybe just rsync it every night or something
There's a thread you and others are pulling on here, and we need to pull it. Hosting doesn't have to be the domain of the big vendors anymore.
toomuchtodo 22 hours ago [-]
That was my intent, pull the thread.
torginus 21 hours ago [-]
Well ackchually.. I get that large scale systems pose their own challenges on their own, but it also matters what's the smallest isolable unit.
What I mean by this is a CDN consists of nodes that are horizontally replicable and don't really talk to each other, and thus are easy to run even at scale.
In contrast, something like a bank or social media isn't really reducible - every user needs to be able to interact with every other user in a consistent manner.
So running a midsize bank's backend which processes 10m transactions per day, might be as if not more complex (all consistent, repeatable, and must never fail), that having a product which is a 10-10k org's IT infra replicated a thousand times.
And yes, lots of people have worked at banks and other fintech companies of this scale, including me.
I am not an expert, as I never worked on the 'core' systems but I know folks who did, and everyone told me there's an arcane database monolith that sits at the heart of these, very expensive and exotic big box SW & HW (at least for us unwashed rubes used to EC2 instances)
r3trohack3r 20 hours ago [-]
> What I mean by this is a CDN consists of nodes that are horizontally replicable and don't really talk to each other, and thus are easy to run even at scale.
This is only true if you exclude problems like “finding a CDN node from the device,” “managing congestion,” etc. as part of the problem statement
supriyo-biswas 18 hours ago [-]
> What I mean by this is a CDN consists of nodes that are horizontally replicable and don't really talk to each other, and thus are easy to run even at scale.
They do though! They mostly try to avoid it since hitting the network to serve any kind of latency would unacceptably increase latency, but you wildly underestimated the amount of complexity there is to running a CDN.
torginus 17 hours ago [-]
I probably underestimated the complexity and I didn't mean to knock on CDNs - I just wanted to say that not all distributed systems have equal complexity, and some require essentially almost serializable transactions, while others are fine with small channels of eventual consistency
dastbe 10 hours ago [-]
the complexity is in different places. For these systems the sheer scale of throughput makes reasoning about them challenging.
the OP mentioned 1B requests/day, where there are systems handling 1B requests a second.
kennethops 21 hours ago [-]
As a previous SRE at Cloudflare, I'll never shit-talk fellow SREs at big companies.
The level of scale and complexity a big tech SRE has to deal with on a constant day-to-day is a very imbalanced proposition. A lot of people, in my experience, are not fully comprehending.
You have to be a jack-of-all-trades and a master of all.
codeulike 18 hours ago [-]
Yes, the thing Salesforce are good at, and is little understood here, is that they've kept their platform online for 27 years so far. Its constantly evolving, three upgrades per year, but changes that require customers to change their customisations are rare, and when they happen they are communicated at least a year in advance. Approx 150,000 tenants, all with different configurations and some so heavily customized that they are effectively unique apps. Salesforce keeps them all online and evolving. In those 27 years there hasnt been a 'lets trash this and rewrite from scratch' and there hasn't been a 'you must migrate your data to our new platform, we're closing the old one'. They've just evolved it while running. They must have got some things very right in the original architecture to be able to do that.
One thing that I find interesting is that they launched their platform language Apex (a sortof subset of Java) in 2007 when TDD was the hot new thing, so TDD is baked into the platform - your Apex code must have at least 75% test coverage, and the tests must pass, before you are allowed to deploy to prod.
They leverage that test coverage when they are upgrading the platform - they have an internal process called The Hammer where they run all customer-created tests against customers own unique configs on the current platform and then again on the next version of the platform to see if any customer tests are being broken. Look it up, its really interesting.
Cthulhu_ 9 hours ago [-]
> They must have got some things very right in the original architecture to be able to do that.
I'd argue they (and many larger, older, established etc systems) may not have, but it's part of how it works so while it may not be the best it's the one that is working right now and earning them money - working (and earning) software always trumps correctness etc, in practice.
8 hours ago [-]
nightski 21 hours ago [-]
For me personally the snark isn't because their SRE team is incompotent. It's because software that tries to be everything to everyone is inherently terrible. It's not fun to use for the users, and so many compromises need to be made on the technology side to make that happen that it ends up just being crap all around. This includes Salesforce, SAP, Dynamics, any platforms like that which scale many industries.
Flexibility and abstraction come at a high cost. It doesn't really matter though, world domination at all costs is the name of the game.
Cthulhu_ 9 hours ago [-]
Thing is, if one piece or suite of software doesn't try to be everything (in a fairly consistent way), businesses that need certain functionality will end up with various different tools from different suppliers, which is at the very least just as complex and expensive to manage, and in practice more expensive.
elzbardico 16 hours ago [-]
AI changes this equation fundamentally, in a way that VC and SaaS founders still haven't realized.
Nobody ever liked having to change their business, their workflows, or ducting taping a customization in a SaaS, they did because as we moved from centralized mainframe apps, to PC client-server and then Web based distributed apps, it became increasingly more cost-effective to suffer with a generic, one-size-fits-all SaaS than building at home.
AI coding changes this a lot.
Cthulhu_ 9 hours ago [-]
I'm going to give this a "maybe"; the challenge with large scale software isn't in authoring new code or whatever, it's in managing complexity.
I think AIs / agents (and more importantly how we are learning to use them effectively) may help in that regard, but only if they are able to manage that complexity. This'll depend on context window sizes, their ability to explore a codebase, and how well their operators can provide relevant information.
But that's only what they can consume (so codebase, documentation, etc), on top of that are the people that work with / for these systems for decades and who know a lot about things outside of what's written down.
gibsonf1 24 hours ago [-]
Hmm, could the use of genAI have anything to do with this failure and the inability to quickly fix it?
Jach 23 hours ago [-]
It's not impossible, but Salesforce has had big outages before LLMs. For a disruption that began at 1am pacific, the response time isn't that bad. 3 hours total to give up on restarts, 4 hours total to validate a quick fix and begin rollout, and the rest of the time since has been waiting for the rollout + addressing subsets of instances that had some issues with restarting+the quick fix. It's nearly 9am pacific now, so Dreamforce is saved~ (It's Dreamforce week this week. Most devs are either focused on that or on soft-vacation / working on lower priority non-feature-work items, it's surprising anything would be updated to production this week that could do this.) The architecture and approval process of everything there has long been setup so that things can't be changed quickly.
karagenit 22 hours ago [-]
The bug was from 2009, so probably not :)
trebligdivad 24 hours ago [-]
What I'm curious about is why it is a single-PaaS; I'd have expected Salesforce to have the customers quite isolated so the chance of bringing down multiple customers at once was much smaller.
stmw 20 hours ago [-]
The customers are quite isolated, but it doesn't mean that some services or errors do not propagate. In public cloud terms, think back on some AWS or Azure or even Gmail outages - you probably wouldn't even hear about them if it didn't affect millions of users at once, across security and availability boundaries.
Cthulhu_ 9 hours ago [-]
I think it's easy to underestimate how many smaller outages there are in any period of time but which do not affect anyone or only a small number of people due to all the mitigations, due diligence that developers and SREs do, and the self-repairing nature of modern systems.
kakwa_ 21 hours ago [-]
You still have fleet wide management which can cause issues. Plus there are always a few core services like queues, authz, presentation layer.
Also, mono-tenant architectures is no golden bullet either. Such architecture (often coming from a formerly on-premise product that was SaaS-ified) can easily become hell to operate as it multiplies the integration points (DB parameters, URLs, allowlists, etc).
It's also quite wasteful in terms of resource utilization and hosting costs.
prpl 23 hours ago [-]
a few things are global (login service) to some extent, just as AWS places several such things in us-east-1
mulmen 1 days ago [-]
I honestly don’t get the snark. The status page has:
Seemingly meaningful IDs
Search
Region filter
Email update signup
Predictable URLs for instance status so they can be deep linked in runbooks
What appears to be the actual live instance status.
What appears to be the actual live service status in each instance.
An update log with frequent detailed updates.
troyvit 23 hours ago [-]
I despise Salesforce, but when I landed on this page I was like, huh. wow. honesty. Looks at GitHub
So yeah you're exactly right, the snark is not deserved if you ask me, and I'm 82% snark.
dominotw 1 days ago [-]
[flagged]
shimman 1 days ago [-]
This is the case for every single B2B saas product. This is like the "bar is rolling on the floor" level of competence required. Please have higher standards for paid products.
eastbound 1 days ago [-]
What do you think of Atlassian?
shimman 23 hours ago [-]
Terrible, I'd argue the vast majority of modern big tech offerings are extremely poor quality where the need for surveillance in the form of constant monitoring/advertising metrics deliberately makes these types of services more costly to maintain and repair over time.
Sure there are like 3 or 5 decent services out there (like S3) but the vast majority are over engineered to be user hostile while extracting out whatever resources they can from their customers.
raffraffraff 1 days ago [-]
Have you tried turning it off and then on again?
> We're no longer pursuing restarts as a path to remediation.
Oh you have
cube00 1 days ago [-]
Kind of surprised they admit they're going to try restarting and see what happens. I'm sure it happens everywhere but nobody admits it.
> We've attempted a rolling restart on one of the impacted instances to see if that resolves the issue.
At least it didn't fix the problem so they can actually start finding the real cause.
> We're no longer pursuing restarts as a path to remediation.
Why isn't the AI they sell telling them what's wrong? Why do they need to take shots in the dark to "see if that resolves the issue"?
swatcoder 1 days ago [-]
I don't know, that reads exactly like an AI troubleshooter working through a plan without the implicit contextual understanding an experienced human might bring to either the actions or the communications.
"Oops, we forgot to tell it that this is the hyperscaled Salesforce production environment and that its choices need to project competence and consider brand embarrassment. WILLFIX"
jakevoytko 1 days ago [-]
In my experience it’s a safe way to do something useful while everyone is getting their bearings. It immediately partitions the situation space between being persisted or systemic vs local or caused by long-running processes. Plus everyone’s going to ask if you’ve tried that already, so you might as well get it out of the way if it makes any amount of sense
vrosas 1 days ago [-]
Followed quickly by "Redeploying with more log lines", the next logical step.
SoftTalker 22 hours ago [-]
In my experience this is very common on linux hosts. Not that you would do a full system reboot as a generic first attempt (this was much more common when I worked with Windows) but restarting a wedged or misbehaving service/daemon is a pretty common thing.
_joel 24 hours ago [-]
Been many an MIR that I've seen where after initial assessment, the next log entry was "service restart attempted"
b112 1 days ago [-]
Restart should be a very last emergency step, as if it works, a restart often might wipe out evidence of why.
So hopefully it's not done often.
wookmaster 3 hours ago [-]
The requirement is to get customers out of impact as #1 priority. If there's suspicions around memory/thread states a restart makes a lot of sense. Digging through logs and flight records takes a lot of time, customers are losing business in that time. If you're afraid to restart your service you need to work on your telemetry.
CoffeeOnWrite 22 hours ago [-]
I wouldn't say so, rather you need to balance recovery time and evidence preservation. A good incident manager will give the service owning team a chance or two to debug, but not let them fall into the trap of needing to understand the problem fully before attempt a clumsy potential fix. And of course will take into account the total business impact of the ongoing disruption and the known and unknown risks of the proposed clumsy fix (it could make things worse).
b112 21 hours ago [-]
Yes, that's why it's the last emergency step. We're not disagreeing.
ocdtrekkie 18 hours ago [-]
Rebooting is the first step: If it fixes it you don't have a problem. If it doesn't, you know more about the problem.
It's a joke, but like, after nearly two decades of engineering I have something break on me, due to updates. I figure the updates broke it, call the vendor, and they go... did you try rebooting it again?
Rebooting it a second time fixed it.
andrewinardeer 1 days ago [-]
"Yeah, I'm with Rob. Just let's reboot and see what happens"
chihuahua 1 days ago [-]
"If that doesn't work, clear the cache and reboot again."
cube00 23 hours ago [-]
[dead]
1 days ago [-]
reaperducer 21 hours ago [-]
Have you tried turning it off and then on again?
Flashbacks to "Thank you for calling Three-Ten-DELL. Have you tried turning it of and turning it back on again?"
mergy 1 days ago [-]
Unplanned outage timing is never good but this is really not good.
Say what you want about the product and leadership... they do throw a good conference.
ramesh31 1 days ago [-]
Probably not a coincidence
dehrmann 22 hours ago [-]
Most places at this scale have code freezes in place well before conferences. The most likely issues are some launch couldn't handle the scale or periodic deployments have been saving them from some sort of long-standing leak bug, and pausing going into Dreamforce meant some service hasn't been restarted in a week. Historically, Salesforce sharded by customer, so that goes against both of these, unless it's in a routing layer.
SaucyWrong 18 hours ago [-]
I'm an ex-Salesforce, and yes, at the time I left, there was a huge change freeze surrounding Dreamforce. Unless a demo of an announced feature was coming in really at the buzzer, change velocity would have been low since a few weeks ago. I worked in a sub-cloud, so I can't even speculate as to the reason for the failure.
Something I wonder about is whether SRE responses were delayed due to having to be emergency-change-approved because Dreamforce was on. I don't recall a global outage ever occurring during a change freeze when I worked there, so /shrug.
karmakaze 24 hours ago [-]
Yup. Could be that everyone was rushing to get all their products and demos ready leading up to it.
gregw2 18 hours ago [-]
AWS has a history of this before re:Invents (its annual big conference) in my experience.
They've gotten a little better in recent years.
24 hours ago [-]
minimeow 1 days ago [-]
This is what happens when more than half the company is away attending the Salesforce cult-indoctrination stuff while spending all their bandwidth making customers/partners feel good.... The stuff that matters to keep the lights on gets overlooked.
gigatree 1 days ago [-]
Arguably, making customers and partners feel good is the more important part of the business
walt_grata 1 days ago [-]
They wont feel good if the product they pay for doesnt work
chaboud 1 days ago [-]
From what I can tell, the business model is "it doesn't work, but you can pay folks exorbitant fees to 'customize' it for you"...
Also see: Oracle
23 hours ago [-]
paimapi 1 days ago [-]
arguably, this is what sales cares about and the half-measures taken to tackle what must be the Mt. Everest of tech debt at Salesforce is what leads to large, systemically degraded customer trust in products that keep shipping bugs
orochimaaru 1 days ago [-]
I don’t think engineering and SRE of the organizing company are ever invited to those events. They’re mainly for marketing and sales (which includes solution architects).
SaucyWrong 18 hours ago [-]
I'm an ex-Salesforce engineer and I can attest that very few line-level engineers are ever invited to attend Dreamforce. This would have been an ordinary day at the office for the SRE team.
dominotw 1 days ago [-]
this event is only for customers. its not a company event.
fhub 1 days ago [-]
Cause: Legacy Salesforce login service got into a resource-exhaustion cascade.
Fix: Rolling some unspecified fix they proved in testing out over the fleet seemingly very slowly (After their earlier attempts to roll something out faster failed).
Coincidentally, Salesforce just laid off the senior engineering lead on their shared login service team who 2 months ago was describing the practices his team follows to keep 99.95% uptime on that shared login service (this guy: https://www.linkedin.com/posts/patricktprice_engineeringmana...)
Layoffs are usually short-term money-savers at long-term cost, but hot damn this is the shortest short term I've ever seen.
dd8601fn 1 days ago [-]
I wonder if “legacy login” is the shared login gateway.
It’s optional but everyone uses it. And it was flaky for an hour or so, like two months ago.
kgermino 20 hours ago [-]
And while I haven’t actually done my homework, the push to email-based login seems like it’s really pushing people to start using their own custom domains for logging in
cyberpunk 1 days ago [-]
That status page is the most salesforce thing ever.
Scroll down. >_<
sunrunner 1 days ago [-]
Wow, it's almost as long as the Every UUID V4 or Every Floating Point Number pages.
dgellow 1 days ago [-]
I see tabs with a spinner loading infinitely, which is indeed very salesforce like
troupo 1 days ago [-]
It lists all instances, you can drill down into each one of them, see which services are affected, and each affected service pops up the incident timeline?
Isn't it actually amazing, and not "the most salesforce thing ever"?
santoshalper 1 days ago [-]
At least they're consistent about UX. My only complaint is that it needs more tabs.
chrisjj 1 days ago [-]
Yup. No mention of outage. Even drilling down gets nothing more than "Service Disruption".
electroweak 1 days ago [-]
At this point OpenAI really ought to let us know when they're testing again.
DoesntMatter22 24 hours ago [-]
My thoughts exactly.
AmazingTurtle 1 days ago [-]
You better bet someone started their agents with a prompt "Make a salesforce clone but with 100% uptime"
Bluestein 1 days ago [-]
"Make a salesforce clone but with 100% uptime"
⎿ You've hit your session limit · resets 2:53am (48°52.6′S, 123°23.6′W Etc/GMT+8)
/upgrade to increase your usage limit.
Finally! You win the Internet today good sir! (At least as far as I am concerned :)
chasd00 1 days ago [-]
You can click on any of the instances and then the service that is down to read the updates. It’s not 100% clear but some sort of issue with a “legacy login service”. The latest updates say a fix is rolling out.
arionmiles 1 days ago [-]
This is certainly a unique status page.
adamw2k 1 days ago [-]
Perfect timing with Dreamforce this week.
reddalo 1 days ago [-]
I can't understand how such a huge company can have such a lousy UX.
WinstonSmith84 24 hours ago [-]
Peak Tech Salesforce was 2010 +/- 2 years - i.e. after Visualforce and before Aura era. It used to be a developer oriented platform and it became shiny/flashy garbage eventually. But all these shiny things allowed them to get a large market cap with very brilliant sales people, it's hard to deny.
dd8601fn 23 hours ago [-]
VF and Aura overlapped. Aura was just a bad start and janky. We sometimes just did React instead, for a while.
LWC is worlds better. And the local tooling with the cli and VSCode extensions is miles better than the old Eclipse/Sublime FMT days.
WinstonSmith84 22 hours ago [-]
Aura existed simply because Salesforce thought to be smarter than open source, well. Can't deny it though: Salesforce engineers were great on the backend, but frontend dev has never been their thing. Back then there was Angular 1 which was miles ahead. React was released shortly after Aura itself, so to say. LWC is what Aura shall have been 10+ years ago. And that ties back to what OP wrote: awful UX extremely slow bloated with JS.
Now, don't talk me about VSCode Extensions. This is the perfect example of an awful dev experience. apex-jorje-lsp.jar with a JVM to parse Apex taking GB of memories, extensions taking dozens of seconds to load (when they load) ... In fact, the only decent LSP is aer, a simple decently working Go binary rather than the monster Salesforce shipped. The one good tooling Salesforce built in the last 15 years is, to some extent, the SF CLI - which came after the `force` CLI from the same guys who built `aer`, anyway. And nowadays, people can use that with their preferred editor from Zed to Vim with shortcuts from built upon the SF CLI.
So no, Salesforce didn't do great with tooling, they just did the bare minimum waiting on the (small) community to give them the right ideas.
dd8601fn 8 hours ago [-]
Angular and React didn’t solve the same requirements, they were just somewhat workable stand-ins while Aura was jank.
And the tooling is heavy, but I’ll repeat, it’s all miles better than it was back then.
We traded janky and sparse for a full suite of pretty great tooling that can be a resource hog.
Remembering the relative simplicity of the old stuff as “better” overall is just rose-tinted glasses.
That's not what I meant. Salesforce had a choice to use (and support) battle tested frameworks, and decided instead to build their own one.
abeyer 18 hours ago [-]
To be fair, they came in with some requirements that I don't think anyone else had at the time, and even today aren't in any mainstream ones afaik.
I think the biggest difference was around providing security and stability barriers between front-end components on the same page, with the intent of allowing you to compose a page that contains your own components and those of other third party applications you've installed with guarantees about how they can (and can't) interact. Not sure they couldn't have tacked that onto another framework, but it comes with enough trade-offs and compromises that I'm not sure anyone else would have wanted to upstream it, so they would have been forking something anyway.
Aura wasn't much fun to work with, was never really feature complete, and not advocating for it... but it actually kind of made sense if you thought about front end with the context of how salesforce did security and multitenancy in mind.
prettychill 19 hours ago [-]
I think around 2012 there was not a lot of options for an enterprise rally around. Aura was designed and built in the same timeframe as react/angular/ember/etc iirc.
ncr100 1 days ago [-]
Is like Microsoft Windows in that way?
reddalo 8 hours ago [-]
Yes, maybe you're right. Windows used to be good in the 9x era. Then it all went downhill.
danjc 1 days ago [-]
It's dns isn't it
whatthesmack 19 hours ago [-]
The last time it was DNS at Salesforce, the poor guy got unceremoniously canned. So there's some incentive for it to... not be DNS.
bearjaws 1 days ago [-]
Feel like it has to be for all of this to go down at the same time.
Bluestein 1 days ago [-]
It's always DNS.-
ajross 1 days ago [-]
More like 70% human-configured DNS, 25% human-configured routing configuration, 5% interesting software bug.
Bluestein 1 days ago [-]
Entirely correct.-
(Nowadays any of those need to fit in an "agent dropped all tables. Apologized" moment.-)
ghusto 1 days ago [-]
No, IPv6
dogas 1 days ago [-]
I was scrolling to see this comment, haha
afr0ck 1 days ago [-]
Classic
alienbaby 1 days ago [-]
Haven't they got some kind of new fancy ai interface they can use to fix it?
cmiles8 24 hours ago [-]
Salesforce has turned into the monolithic messy bloatware it set out to replace.
Please VCs stop with the AI FOMO and find a few good startups to just go destroy Salesforce and give folks a simple inexpensive replacement.
dd8601fn 23 hours ago [-]
There are 10,000 “simple inexpensive replacements”. Always have been. If you need a glorified three object contact list… you shouldn’t buy Salesforce.
Those just, obviously, can’t do almost any of the forty million serious things that Salesforce does, and real businesses do need.
stefanfisk 22 hours ago [-]
This is an honest question: which serious things?
bobsmooth 18 hours ago [-]
My boss recently added a custom integration to allow us to log hours on jobs. It includes GPS coordinates so we have proof of our whereabouts.
stefanfisk 10 hours ago [-]
But that’s not a salesforce feature but rather custom software. You could bolt that onto most systems.
bobsmooth 8 hours ago [-]
Sure, but it's already integrated in the Field Service app we already use.
sschueller 24 hours ago [-]
As someone who has never used Salesforce nor hubspot. What are the core features that has people using these services? Is it just the integration between all the different areas where customer data lives?
cmiles8 24 hours ago [-]
It’s like a lot of these big SaaS platforms that lock people in to big contracts. They promise the world with all sorts of features and integrations but most customers don’t use that and get locked into buying a very expensive solution to a very simple problem (in this case tracking sales pipelines).
Where customers do use the features and integrations it’s often a giant mess that needs a whole separate ecosystem of consultants and “partners” to get the thing working and maintaining it.
kakwa_ 5 hours ago [-]
Integration and reconciliation between arbitrary data sources (from SFTP+CSV files to connecting to data warehouse like snowflakes, potentially through vpn tunnels).
Data querying for creating marketing campaign audiences or generating some reporting reporting.
Little development for landing pages, forms, or custom integrations.
And many marketing stuff like AB testing/fatigue, etc
IneffablePigeon 22 hours ago [-]
They're an incredible marketing machine.
But mostly it's a mix of the integration network effects you mention, cost of reimplementation if you want to leave, and the good old "nobody gets fired for buying IBM" dynamic.
All the prompt engineers clogging up San Francisco for dreamforce need to get back to the office.
ssk42 1 days ago [-]
My gut instinct is that this is about when all of their on prem servers were EOL and their /public cloud solution was required. This must have had something to do with that
BoorishBears 1 days ago [-]
My gut instinct is that this is about Dreamforce with the rickshaws and whatnot
bgro 1 days ago [-]
Leetcode developers win again
voidUpdate 1 days ago [-]
Wow, the intern must have tripped over a very big power cable this time
cmiles8 1 days ago [-]
Dreamforce, wake up. You’ve overslept!
mariopt 1 days ago [-]
Could it be people doing Claude/GPT automations and they just can't handle it?
mulmen 1 days ago [-]
Could be. With a big conference going on and the updates mentioning resource exhaustion it could be a bunch of people doing demos, AI driven or not. Basically slashdotting themselves.
lrvick 1 days ago [-]
Seems it is back up now. Damn.
jason_zig 23 hours ago [-]
I was just working on a zigpoll integration and thought it was me...
sidcool 1 days ago [-]
ClaudeForce in action!
Havoc 1 days ago [-]
So I guess today everyone gets actual work done
theshrike79 1 days ago [-]
Was it DNS? Any guesses? :)
dude250711 1 days ago [-]
[flagged]
geerlingguy 1 days ago [-]
They let their agentic AI take over system maintenance at Dreamforce yesterday like they were pushing in the talks
/s (partly)
jtrn 1 days ago [-]
And nothing of value was lost. God, I hate everything about Salesforce. Sometimes I have to integrate against their services, and it is always a pain, not to mention what the core project actually is: optimization of marketing and spam.
dd8601fn 24 hours ago [-]
Sure, if you’re struggling with weird custom endpoint behavior, or custom triggers and validation, that’s on your admins and devs.
But if you’re having trouble with the standard rest or bulk apis, that’s 100% on you.
jtrn 4 hours ago [-]
If you think integration with Salesforce = REST API then you have no idea what you are talking about.
felineflock 23 hours ago [-]
What is painful about a plain REST API?
jtrn 4 hours ago [-]
A normal sensible REST api ? Nothing.
Salesforce API?
Sandboxes drift from production in both metadata and data. Refreshes change record ids.
API access is gated by edition. Professional edition has none by default. Salesforce sells MuleSoft as the fix to their own shit architecture.
Custom objects and fields mean that connecting successfully doesn’t establish what represents a customer, subscription or completed sale. Even standard objects can be used differently between organizations.
Metadata deployment is SOAP-based and slow. Dependencies between components break deploys and is massive pain to debug.
And notice that I said integration, not API integration. There’s a shit storm of terrible way to integrate beyond just simple API. Have fun with Bulk 1 or 2, Composite, Streaming, Platform Events, Change Data Capture, Tooling, GraphQL. An my most hated item, anything web based.
I seriously want to talk to someone that have a good time integrating against Salesforce. Maybe I will have my mind blown as to how easy it could be, but I suspect I will have the same experience I have every time I critical scrutinize and such situation: it’s just as terrible as it seems and the person claiming it’s easy is hands down producing close to nothing of value.
chasd00 1 days ago [-]
not really defending salesforce but OAuth+REST is a pain? Pretty plain vanilla in terms of integration requirements.
sergiotapia 23 hours ago [-]
RPC errors in my Dunkin app and my wife's Poshmark app. Is it an infrastructure problem?
dboreham 1 days ago [-]
At least now we can figure out what Salesforce does.
sunrunner 1 days ago [-]
"And in tonight's news, the worldwide CRM solution Salesforce had a global outage affecting one hundred percent of its customer base. We interviewed users of the service to find out the scope of the impact. Everyone agreed that they were impacted, but strangely, nobody could describe _in what way_ they were affected."
zzzeek 1 days ago [-]
I'm sure the cause of this outage will not be connected to vibe coding in any way
sparkling 1 days ago [-]
Remind me please, what are folks currently paying per seat for this glorified CRUD app?
warmedcookie 1 days ago [-]
It's this thing called golf course driven development
electroweak 1 days ago [-]
Permanent cache for that one.
abeyer 18 hours ago [-]
There is a published sticker price, but I don't think anyone outside of tiny installs with just a few users pays that... it's very much the enterprise sales model of "let's schedule a call and talk about it." They also for many years would let previously negotiated prices stand during renewals even when sticker price went up, so older orgs often have the "same" license being charged at many different prices based on when it was first purchased and how the negotiations went at the time.
raverbashing 1 days ago [-]
If I had been just out of Uni I would think this is edgy
Now I'm just glad I'm not responsible for this fire
aeneas_ory 1 days ago [-]
It’s akways DNS or login
wronex 24 hours ago [-]
Does it mean I won’t get any AI generated slop spam for a few hours?
jackdecker 1 days ago [-]
Ah yes. Exactly what a status page should look like: an endless list of random ID’s that don’t mean anything and no information whatsoever
At least salesforce is consistent with their design language
chasd00 1 days ago [-]
> random ID’s that don’t mean anything and no information whatsoever
If you use salesforce you know what all of that stuff means. Just click on one, it’s not rocket surgery.
jackdecker 1 days ago [-]
Was really just poking fun at them - AWS’s status page isn’t much better
1 days ago [-]
dd8601fn 1 days ago [-]
Random Ids? If you mean the “USA324” ones, those are pods. If you’re a customer you know which one(s) you care about.
xboxnolifes 22 hours ago [-]
This might actually be one of the most useful status pages ive seen. Its not just random green tick marks representing the entire service that only change yellow when someone gives and admits that 5 hours of bad service is an outage.
BoorishBears 1 days ago [-]
Looks like a region list to me, maybe just with a lot of regions
1 days ago [-]
1 days ago [-]
themgt 1 days ago [-]
Never before in the history of global compute outages was so little lost by so many down servers, whose purpose was known to so few.
elzbardico 1 days ago [-]
Kind of ironic. Salesforce is basically one of the major spiritual grandfathers of Slop. It is not uncommon in production systems to find that objects like Contact and Account have hundreds of custom fields. Sometimes, you find out that several of them have the same meaning and semantics, but were used at different times. Digging out you discover that some Marketing guy that used to work at the company did some task in a certain way that was lost when he was gone, and then a few months later his substitute had the same need and went ahead and created the same field with a slightly different name.
Doing data engineering work with Salesforce data is an exercise on archeology, psychology and organizational politics.
Slop is basically the ontological and teleological philosophy behind Salesforce very existence. Despite the official discourse that the "No Software" meant no infrastructure, no toil with updates and configuration, the subtext as intended for executives was very clear: "No need for you to be blocked by those pricks from engineering and their stupid, bureaucratic and gatekeeping rules".
"No software" was a call-to-arms to a certain subset of managers that were radicalized by Nicholas Carr's 2023 HBR article "IT Doesn't matter". It doesn't matter that Carr was a journalist and a writer with a masters in English that has never ever run even a small bodega, or has never managed an IT department. Anti-intellectualism and the abundance of capital brought in by the petrodollar that allowed the US government to run deficits year by year while exporting the ensuing inflationary effects to rest of world, would ensure that this message would ressonate and then even be amplified during the years of ZIRP and the Baillouts. Play fast and loose, first come, first served, a rising tide rises all boats and all that jazz. Wall Street favors bold, and the heck with the long term! This quarter will only live once!
Frankly, this is just poetic justice: Kill by slop, be killed by slop.
gregw2 24 hours ago [-]
Typo: Carr wrote that in 2003, not 2023 for those of you who missed the foolishness and insane wreckage that article caused.
The lost business value and competitiveness caused by outsourcing IT overseas to unmotivated parties under Carr's premise is hard to put your finger on but I have seen the aftermath and it's pretty massive.
elzbardico 22 hours ago [-]
Thanks for catching the typo!
abeyer 18 hours ago [-]
> Doing data engineering work with Salesforce data is an exercise on archeology, psychology and organizational politics.
What business focused systems is this _not_ true for? Doubly so for such systems that encourage non-developer users to customize the database schema?
Not saying that's always a great choice, but they're far from the only ones to have made it, enterprise customers love to buy it, and Salesforce is actually one of the better ones I've seen to deal with from a technical perspective. There is _far, far_ worse out there. Most of this comes down to a business and design problem... your entire first paragraph reads as a broken process that no system is going to fix.
schnevets 1 days ago [-]
Something tells me Troy the Salesforce Admin/BD Analyst did not cause the SAAS infrastructure to go down.
And I think you're confusing crud with slop.
isityettime 24 hours ago [-]
I've never worked somewhere that had a Salesforce integration which wasn't an eternal disaster. Have you?
Why is every company's Salesforce team absolute bottom of the barrel developers with super high churn, no responsibility, and little competency?
Something about the product and its positioning attracts catastrophe. That's what GP is talking about.
abeyer 18 hours ago [-]
It sounds like you've worked places that have made a staffing decision and paid the price. It happens, but it's far from universal.
Salesforce is not simple. It's wildly, overly complex. It's amazing it has any 9's at all and not 8's or 7's. Salesforce offers three 9's, which allows for 43 minutes downtime per month. The current outage is at 8 hours (and counting) so Salesforce is now at 98.9% uptime for the month - there's an "8" in there now. Not good, but considering the complexity of Salesforce, it's still kind of amazing.
It turns out business environments are wildly overly complex.
I remember Cisco before iOS used to have hundreds of branches for their router, one branch for each major customer that was demanding specific features. It was unmanageable, but that's what you needed to do to win those "enterprise customers".
It also turns out customers aren't very good at articulating their needs and putting them into a cohesive vision of the product. But they sure have specific demands to get stuff in. I'm not blaming the customer, this is just how this world works -- All of the "enterprise software" apps are extremely complex with hidden knobs and weird behavior that was pushed in by a customer twenty years ago all over the place.
Ouch. I heard of a company in my home town (small B2B service provider) doing something similar - they paid well but I didn't think it was worth it.
But the GP is essentially correct - there is a 2% of salesforce that could be built run and keep 80% of salesforce users happy. Except that you could not charge enough to be able to advertise on F1 cars and take SVPs out to dinner.
So you could not actually make 80% of them happy - they would ever buy it.
Yes, you largely do - they’re the commits that get rushed to, and through.
This take that showstopping technical debt is unavoidable is very new, and will age like milk.
No it's not. The push and pull between shipping and paying down technical debt is as old as there's been software to sell. Sales has been selling features that don't exist quite yet ever since they've been talking to customers, and engineering has been pushing back on implementing them yesterday since there's been features to implement. Showstopping technical debt is merely a side effect of who wins that argument in a given org.
My point is the technical debt is stopping the show way more often.
Not “this never existed before selling”,
but “we never had the team in place who could do this right in the first place”.
You can quibble about who is responsible, but the fact remains.
I think you are missing the point. When I state my Exchange server is more reliable than Exchange Online, I don't think I'm a better engineer. I recognize Microsoft has harder problems to solve than I do. I think building overengineered, oversized SaaS environments is introducing extreme risk. It's an inherent flaw of the current approach.
Smaller is, in fact, better, because it's easier to operate reliably.
I think people forget that those large environments are there for a reason. To make sure the service stays up in the face of problems outside your own control.
https://hn.algolia.com/?dateRange=all&page=0&prefix=false&qu...
My entire point is that you have no redundancy in your system and you also aren’t big enough to have any pull with the vendors who can fix these types of outages so you’re basically at the mercy of your providers with no recourse.
That’s why these systems are built the way they are.
And generally four nines is considered the gold standard these days. I can tell you for sure that both Netflix and Ebay would lose money anytime they drop below four nines because I have at some point been responsible for both. You’re correct that Reddit has a lot more leeway and outage time before they start losing money but not that much leeway.
Built how? Because I can state with confidence that I have cleaned up a ton of failed upgrades/zombie terraform deploys of these serverless kubernetes wonders that followed every best practice under the sun, and these things are not really considered even moderately reliable (as designed by imperfect mortals under real world conditions), meanwhile professionally, people who stand to lose a lot of money should their systems go down generally operate systems whose architectures were designed decades ago, are generally horizontally scaled monoliths, and are extremely conservative in software choice.
Also downtime often is no biggie, as long as it's planned and or don't lose (too much) customer critical data.
Like nodobody cares if your test db cluster goes down for the weekend. We even shut down our db instances to save money.
Even if a solar flare takes out an entire continent or two I think it's safe to say that the bittorrent network will still be running in some form. Can you be so certain about any given SaaS product?
"What are you optimizing for?" is always an important question, as is "The Five Whys."
Over multiple ISPs, so far internet outages for more then a few minutes is very rare (though the few minutes would make me not want to host something requiring high availability; and a cut cable is really annoying because there simply is no quick fix), power outages even rares, I experienced 3 in 40 years, and the longest was 6 hours.
Ignoring for now how you are synchronizing the database and filesystem, and how doing so may well result in your duplicate experiencing the same failure as the original, you can maybe recover from a small class of availability issues that could knock you out of an SLA.
But that assumes you can get online and can fully orchestrate the transition within less than 53 minutes of it starting. Including the time you took to become aware of it. And including the time to diagnose and decide that a switchover would resolve the problem. Including the time it takes for DNS caches to expire and point to the new host. Including the DNS caches which may ignore your TTL. And including all these things again when you switch back.
And assuming, of course, that it doesn't happen again for a whole year.
(I hope you'll agree that the middle east outage is a true outlier)
"Simplicity is the ultimate sophistication." -- Da Vinci
Depending on your time and appetite for tinkering with all of this, it's not hard to imagine a home setup that fails over to a cheap Hetzner or DO VM. A manual failover at the DNS level isn't overly complex, and could be scripted.
Keeping a database in sync between home and the instance might be simple or more complex depending on needs, but would it really be that hard to have Claude help you setup a replicating Postgres server? If your database (or data files) are 1 gigabyte and don't update that often... maybe just rsync it every night or something
There's a thread you and others are pulling on here, and we need to pull it. Hosting doesn't have to be the domain of the big vendors anymore.
What I mean by this is a CDN consists of nodes that are horizontally replicable and don't really talk to each other, and thus are easy to run even at scale.
In contrast, something like a bank or social media isn't really reducible - every user needs to be able to interact with every other user in a consistent manner.
So running a midsize bank's backend which processes 10m transactions per day, might be as if not more complex (all consistent, repeatable, and must never fail), that having a product which is a 10-10k org's IT infra replicated a thousand times.
And yes, lots of people have worked at banks and other fintech companies of this scale, including me.
I am not an expert, as I never worked on the 'core' systems but I know folks who did, and everyone told me there's an arcane database monolith that sits at the heart of these, very expensive and exotic big box SW & HW (at least for us unwashed rubes used to EC2 instances)
This is only true if you exclude problems like “finding a CDN node from the device,” “managing congestion,” etc. as part of the problem statement
They do though! They mostly try to avoid it since hitting the network to serve any kind of latency would unacceptably increase latency, but you wildly underestimated the amount of complexity there is to running a CDN.
the OP mentioned 1B requests/day, where there are systems handling 1B requests a second.
The level of scale and complexity a big tech SRE has to deal with on a constant day-to-day is a very imbalanced proposition. A lot of people, in my experience, are not fully comprehending.
You have to be a jack-of-all-trades and a master of all.
One thing that I find interesting is that they launched their platform language Apex (a sortof subset of Java) in 2007 when TDD was the hot new thing, so TDD is baked into the platform - your Apex code must have at least 75% test coverage, and the tests must pass, before you are allowed to deploy to prod.
They leverage that test coverage when they are upgrading the platform - they have an internal process called The Hammer where they run all customer-created tests against customers own unique configs on the current platform and then again on the next version of the platform to see if any customer tests are being broken. Look it up, its really interesting.
I'd argue they (and many larger, older, established etc systems) may not have, but it's part of how it works so while it may not be the best it's the one that is working right now and earning them money - working (and earning) software always trumps correctness etc, in practice.
Flexibility and abstraction come at a high cost. It doesn't really matter though, world domination at all costs is the name of the game.
AI coding changes this a lot.
I think AIs / agents (and more importantly how we are learning to use them effectively) may help in that regard, but only if they are able to manage that complexity. This'll depend on context window sizes, their ability to explore a codebase, and how well their operators can provide relevant information.
But that's only what they can consume (so codebase, documentation, etc), on top of that are the people that work with / for these systems for decades and who know a lot about things outside of what's written down.
Also, mono-tenant architectures is no golden bullet either. Such architecture (often coming from a formerly on-premise product that was SaaS-ified) can easily become hell to operate as it multiplies the integration points (DB parameters, URLs, allowlists, etc).
It's also quite wasteful in terms of resource utilization and hosting costs.
Seemingly meaningful IDs
Search
Region filter
Email update signup
Predictable URLs for instance status so they can be deep linked in runbooks
What appears to be the actual live instance status.
What appears to be the actual live service status in each instance.
An update log with frequent detailed updates.
So yeah you're exactly right, the snark is not deserved if you ask me, and I'm 82% snark.
Sure there are like 3 or 5 decent services out there (like S3) but the vast majority are over engineered to be user hostile while extracting out whatever resources they can from their customers.
> We're no longer pursuing restarts as a path to remediation.
Oh you have
> We've attempted a rolling restart on one of the impacted instances to see if that resolves the issue.
At least it didn't fix the problem so they can actually start finding the real cause.
> We're no longer pursuing restarts as a path to remediation.
Why isn't the AI they sell telling them what's wrong? Why do they need to take shots in the dark to "see if that resolves the issue"?
"Oops, we forgot to tell it that this is the hyperscaled Salesforce production environment and that its choices need to project competence and consider brand embarrassment. WILLFIX"
So hopefully it's not done often.
It's a joke, but like, after nearly two decades of engineering I have something break on me, due to updates. I figure the updates broke it, call the vendor, and they go... did you try rebooting it again?
Rebooting it a second time fixed it.
Flashbacks to "Thank you for calling Three-Ten-DELL. Have you tried turning it of and turning it back on again?"
https://www.salesforce.com/dreamforce/
Sept 15-17
Something I wonder about is whether SRE responses were delayed due to having to be emergency-change-approved because Dreamforce was on. I don't recall a global outage ever occurring during a change freeze when I worked there, so /shrug.
They've gotten a little better in recent years.
Also see: Oracle
Fix: Rolling some unspecified fix they proved in testing out over the fleet seemingly very slowly (After their earlier attempts to roll something out faster failed).
Details at https://status.salesforce.com/incidents/20004433
Layoffs are usually short-term money-savers at long-term cost, but hot damn this is the shortest short term I've ever seen.
It’s optional but everyone uses it. And it was flaky for an hour or so, like two months ago.
Scroll down. >_<
Isn't it actually amazing, and not "the most salesforce thing ever"?
LWC is worlds better. And the local tooling with the cli and VSCode extensions is miles better than the old Eclipse/Sublime FMT days.
Now, don't talk me about VSCode Extensions. This is the perfect example of an awful dev experience. apex-jorje-lsp.jar with a JVM to parse Apex taking GB of memories, extensions taking dozens of seconds to load (when they load) ... In fact, the only decent LSP is aer, a simple decently working Go binary rather than the monster Salesforce shipped. The one good tooling Salesforce built in the last 15 years is, to some extent, the SF CLI - which came after the `force` CLI from the same guys who built `aer`, anyway. And nowadays, people can use that with their preferred editor from Zed to Vim with shortcuts from built upon the SF CLI.
So no, Salesforce didn't do great with tooling, they just did the bare minimum waiting on the (small) community to give them the right ideas.
And the tooling is heavy, but I’ll repeat, it’s all miles better than it was back then.
We traded janky and sparse for a full suite of pretty great tooling that can be a resource hog.
Remembering the relative simplicity of the old stuff as “better” overall is just rose-tinted glasses.
I think the biggest difference was around providing security and stability barriers between front-end components on the same page, with the intent of allowing you to compose a page that contains your own components and those of other third party applications you've installed with guarantees about how they can (and can't) interact. Not sure they couldn't have tacked that onto another framework, but it comes with enough trade-offs and compromises that I'm not sure anyone else would have wanted to upstream it, so they would have been forking something anyway.
Aura wasn't much fun to work with, was never really feature complete, and not advocating for it... but it actually kind of made sense if you thought about front end with the context of how salesforce did security and multitenancy in mind.
(Nowadays any of those need to fit in an "agent dropped all tables. Apologized" moment.-)
Please VCs stop with the AI FOMO and find a few good startups to just go destroy Salesforce and give folks a simple inexpensive replacement.
Those just, obviously, can’t do almost any of the forty million serious things that Salesforce does, and real businesses do need.
Where customers do use the features and integrations it’s often a giant mess that needs a whole separate ecosystem of consultants and “partners” to get the thing working and maintaining it.
Data querying for creating marketing campaign audiences or generating some reporting reporting.
Little development for landing pages, forms, or custom integrations.
Multiple channels (SMS, push, email) & message templating.
And many marketing stuff like AB testing/fatigue, etc
But mostly it's a mix of the integration network effects you mention, cost of reimplementation if you want to leave, and the good old "nobody gets fired for buying IBM" dynamic.
/s (partly)
But if you’re having trouble with the standard rest or bulk apis, that’s 100% on you.
Salesforce API?
Sandboxes drift from production in both metadata and data. Refreshes change record ids.
API access is gated by edition. Professional edition has none by default. Salesforce sells MuleSoft as the fix to their own shit architecture.
Custom objects and fields mean that connecting successfully doesn’t establish what represents a customer, subscription or completed sale. Even standard objects can be used differently between organizations.
Metadata deployment is SOAP-based and slow. Dependencies between components break deploys and is massive pain to debug.
And notice that I said integration, not API integration. There’s a shit storm of terrible way to integrate beyond just simple API. Have fun with Bulk 1 or 2, Composite, Streaming, Platform Events, Change Data Capture, Tooling, GraphQL. An my most hated item, anything web based.
I seriously want to talk to someone that have a good time integrating against Salesforce. Maybe I will have my mind blown as to how easy it could be, but I suspect I will have the same experience I have every time I critical scrutinize and such situation: it’s just as terrible as it seems and the person claiming it’s easy is hands down producing close to nothing of value.
Now I'm just glad I'm not responsible for this fire
At least salesforce is consistent with their design language
If you use salesforce you know what all of that stuff means. Just click on one, it’s not rocket surgery.
Doing data engineering work with Salesforce data is an exercise on archeology, psychology and organizational politics.
Slop is basically the ontological and teleological philosophy behind Salesforce very existence. Despite the official discourse that the "No Software" meant no infrastructure, no toil with updates and configuration, the subtext as intended for executives was very clear: "No need for you to be blocked by those pricks from engineering and their stupid, bureaucratic and gatekeeping rules".
"No software" was a call-to-arms to a certain subset of managers that were radicalized by Nicholas Carr's 2023 HBR article "IT Doesn't matter". It doesn't matter that Carr was a journalist and a writer with a masters in English that has never ever run even a small bodega, or has never managed an IT department. Anti-intellectualism and the abundance of capital brought in by the petrodollar that allowed the US government to run deficits year by year while exporting the ensuing inflationary effects to rest of world, would ensure that this message would ressonate and then even be amplified during the years of ZIRP and the Baillouts. Play fast and loose, first come, first served, a rising tide rises all boats and all that jazz. Wall Street favors bold, and the heck with the long term! This quarter will only live once!
Frankly, this is just poetic justice: Kill by slop, be killed by slop.
The lost business value and competitiveness caused by outsourcing IT overseas to unmotivated parties under Carr's premise is hard to put your finger on but I have seen the aftermath and it's pretty massive.
What business focused systems is this _not_ true for? Doubly so for such systems that encourage non-developer users to customize the database schema?
Not saying that's always a great choice, but they're far from the only ones to have made it, enterprise customers love to buy it, and Salesforce is actually one of the better ones I've seen to deal with from a technical perspective. There is _far, far_ worse out there. Most of this comes down to a business and design problem... your entire first paragraph reads as a broken process that no system is going to fix.
And I think you're confusing crud with slop.
Why is every company's Salesforce team absolute bottom of the barrel developers with super high churn, no responsibility, and little competency?
Something about the product and its positioning attracts catastrophe. That's what GP is talking about.