> For clean Markdown of any page, append .md to the page URL.
> For a complete documentation index, see https://mailchimp.com/developer/llms.txt.
> For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://mailchimp.com/developer/_mcp/server.

# Mailchimp Engineering

## Reintroducing Mailchimp Transactional for your Email and SMS needs

Mailchimp's transactional email service has long been the go-to engine for reliable, event-triggered communication, whether developers came to know it as Mailchimp Transactional or by its former name, Mandrill\*. Mailchimp's infrastructure sends up to 157 billion emails per month, and Mailchimp Transactional has gained a reputation as a scalable platform that can handle any volume a business demands, from startup to enterprise. Our users know they can count on a strong foundation of 99.99% uptime, an average email delivery rate of over 99%, and a fast median delivery speed of less than 1 second to customers’ inboxes.

But while this mission-critical tool has endured as a reliable, consistent center for one-to-one customer updates, we've been monitoring your feedback about how we can modernize the platform without changing the functionality you rely on. As an engineer myself, it's a balance I understand: Your workflow needs to be as efficient and consistent as possible, even as certain components evolve or improve. Today, we're proud to announce a massive improvement in the core developer experience for Mailchimp Transactional—and share some of the exciting optimizations and additions to the product that you can expect in the coming months.

First, we’ve given the Mailchimp Transactional UI a complete refresh. We’ve modernized the Transactional interface to reflect the modernized experience offered across Mailchimp's other core marketing tools. We see this investment as a crucial step toward a seamless marketer-developer collaboration space. We’ve also enhanced our [documentation](/transactional/docs/fundamentals). We know that clear guidance is critical for faster integration, so we've updated parts of our [API docs](/transactional/api) to include more [real-world code samples](https://github.com/IntuitDeveloper/Transactional-Node-Samples) in major languages, and improved our onboarding guides to help you get to your first API call faster.

Our new API documentation also follows OpenAPI standards, making it easier to integrate our tools into your existing workflows. This approach provides a clear, consistent contract for our API, enabling you to quickly build with confidence.

But this is just the beginning. We're also excited to announce that we’ve added new channel, - [Transactional SMS](/transactional/guides/send-first-sms)\* which allows businesses to send non-promotional, one-to-one text messages triggered by specific customer actions or events, complementing the existing transactional email product. Now available in 10+ countries—the United States, Canada, Australia, United Kingdom, Germany, Austria, Switzerland, the Netherlands, France, Spain, and Ireland.

With this release, you'll be able to see your transactional SMS messages right from your Transactional dashboard, providing a unified view of your omnichannel communications and further solidifying Mailchimp Transactional’s place as Mailchimp’s vital engine for event-triggered messaging. The new Mailchimp Transactional SMS capability leverages Mailchimp’s existing infrastructure and expertise in deliverability to help with compliance and  ensure that critical messages reach the intended recipients effectively. This launch is just the latest development  in Mailchimp’s strategy to invest in and improve its transactional services, building upon the foundation of Mailchimp Transactional (formerly, Mandrill).

By prioritizing a best-in-class developer experience today, we’re laying the groundwork to help free up engineering time for mission-critical work while giving marketers the control they need to move fast. Stay tuned for more details and updates in the transactional space in the coming months.

\*\*\*Disclaimers: \*\*

1. Transactional Email is available as an add-on to Standard and Premium plans, or the legacy Monthly Plan. Pricing varies.
2. Features and functionality vary by plan and certain features may have limited availability. Transactional SMS is available in select countries for Mailchimp users with active SMS and Transactional Email (Mandrill) plans.  Availability is subject to change. See [Mailchimp.com](http://mailchimp.com) for details.

## Mailchimp and WordPress

Image Credit: Devin Wistendahl

Here at Mailchimp Engineering, we’re passionate about empowering our developer community. That’s why we share stories about the triumphs and challenges that our external developers experience while building on Mailchimp.

Recently, we spoke with Aaron Speer, Senior Software Engineer at Gravity Forms. The WordPress form plugin, Gravity Forms, has its own Mailchimp Add-On which enables deep integration between the plugin and Mailchimp. This integration allows users to send data collected through forms on their WordPress websites to their Mailchimp dashboard.

We chatted with Aaron about development at Gravity Forms and about his experience building and maintaining the Mailchimp Add-On.

### What can you do with Gravity Forms?

Gravity Forms is a powerful form builder for WordPress. The plugin allows users to build just about any type of custom data capture form for WordPress websites - from basic contact forms, surveys, and quizzes, to payment forms and advanced solutions for many niche requirements.

The company behind Gravity Forms, Rocketgenius, was founded in 2008, and over the last 14 years the plugin has become renowned in the WordPress community for its advanced features and functionality, robust security and continued reliability.

Aaron told us:

*“It’s almost silly to just call Gravity Forms a form plugin, because at this point it’s more like a user interaction ecosystem. Website owners and operators can use it to interact with their customer base and visitors in an impactful way. They can get detailed information—everything from contact to payment. And in a lot of ways our success has been due to our focus on backwards compatibility.*

*We want that person who installed it in 2012 to be able to keep updating and not worry about it. As a developer it can be tricky to support something that came out almost a decade ago—so much of the web is almost ethereal—but I think it’s paid off. People have confidence that they can install Gravity Forms, have it hang out on their website, and it’s not going to give them issues. I think that’s fantastic.”*

### The origin of Gravity Forms Mailchimp Add-On

One of the key features of Gravity Forms is its extensive ecosystem of integrations. Referred to as add-ons, Gravity Forms has built and maintained a wide assortment of add-ons for numerous email marketing and CRM providers, payment processors, anti-spam tools, and business automation solutions.

In the early days of Gravity Forms, the developers quickly realized that if the plugin was going to collect information, then they should probably make it as easy as possible for customers to use that information. This led to the development of integrations with other platforms, including Mailchimp.

*“It was actually one of the first integrations we made,”* says Aaron. *“Of course, one of the popular uses for signup forms is for newsletters—and when many of us think of newsletters, we immediately think of Mailchimp. Creating an integration was an easy call, so we dove straight in.”*

### Building on Mailchimp: Getting Started

While the Rocketgenius team now has experience creating add-ons, they had more fundamental questions to answer before they could even begin building. They had to determine, first and foremost: “What does connecting Gravity Forms to something else even look like?”

[Learn more about the Gravity Forms Integration](https://mailchimp.com/integrations/gravityforms/)

\_“We came up with a concept we call ‘feeds.’— in an architectural sense, when a form is submitted we take those inputs and feed them to somewhere else. Once we had that in mind, it was pretty straightforward to set up various administration settings pages that were in line with how WordPress dashboards work in general. \_

\_We didn’t want to present a UI to users that was different to everything else there, which would be a frustrating experience, so there wasn’t a ton to do as far as UI or UX goes. We just created an interface where people could look at fields in their forms and choose where the data from each of those fields fed to. \_

*And once we had the Mailchimp side set up, it became relatively trivial to connect those UI dots and give people a way to map from one to the other.”*

### Working with Mailchimp

The Product team at Gravity Forms relies on a number of different development tools. They use PhpStorm, a cross-platform IDE, to quickly set up machines and environments with the other tools they need. \_It’s our Swiss army knife, \_Aaron says.

He also vouches for Insomnia for their RESTful needs:

*“It’s similar to Postman, but newer, and really nice for working with APIs because it’s set up to talk to them exactly how you need it to. I know some companies provide API explorers on their developer pages for previewing what you’re going to see, but I like getting my hands dirty with raw requests in Insomnia.”*

They also used the PHP SDK, although that’s normally something they avoid because of their focus on backwards compatibility. Gravity Forms targets PHP 5.6 as a minimum supported version, so instead of using SDKs that are regularly updated the Product team “hand roll” their own libraries.

*“But, honestly, we’ve never had to do a lot of updates, because Mailchimp’s API is rock-solid and industry standard. Making our own can be a nightmare when there’s a lack of documentation, or the API is nonsensical, but working with Mailchimp was as comfortable as putting on a nice pair of slippers.*

\_I have a lot of experience here—going back a long way, to before the world of RESTful APIs, and before even the idea of shuttling information around the web in a standardized way—but there weren’t any speed bumps. Every request was what I expected, and every response was too. It just made sense. \_

*Our most recent significant integration update came in late 2021, when we switched to using OAuth, and the documentation site had had a major visual overhaul that was very nice as well. It’s not often you see companies put as much design mojo into developer-facing things.*

*If you’ve had negative experiences working with APIs elsewhere, like with some monster like Google or AWS, then you’ll be pleasantly surprised by the Mailchimp API. There are far fewer hoops to jump through, and for developers just trying to get their feet wet with APIs in general Mailchimp is a wonderful one to start with.”*

### What’s next for Gravity Forms and their Mailchimp Add-On?

As the Gravity Forms user base continues to grow, so do the use cases for the plugin. Trusted by non-profits, government organizations, major universities, and big businesses, Gravity Forms looks forward to continuing the support and development of their Mailchimp Add-On - a must have integration that helps ensure the needs of its customers are met across the board.

## Leading Teams and Closing Loops

Lorenzo Gritti

*After two years working at Mailchimp as a software engineer, Soyo Awosika-Olumo recently became a tech lead for the first time on Mailchimp*’s\_ Campaign Builder team. Here, she reflects upon her new responsibilities and experiences:\_

My journey to becoming a tech lead started when I was still an individual contributor (IC) on one of Mailchimp’s website squads. I enjoyed it, but after about a year and a half I wanted a new challenge—and then I saw a position had opened up for a tech lead role on a new project. From speaking to the engineering manager, it sounded like what I wanted to do next with my career. I knew that I didn’t want to become a manager myself, but I did want to be more visible and keep growing as an IC. That’s how I came to join the Campaign Builder team.

Our goal as a team is to build a product that allows our customers to plan, organize, and track their marketing assets and goals in a single place. Mailchimp is like any other tech company in that you start to see the same patterns everywhere in terms of taking data and transferring it, visualizing it, cleaning it up, and using it. Having spent so much time building out services and UI, I wanted to be on the other side—to be the person defining what those services should look like, while also being part of a cross-functional team that I could like and respect, and which supported me in turn.

I also had this long-standing desire to do what I call “closing the loop”—that is, finishing the projects that I start. I’d worked on campaigns before on an entirely separate project. That got shelved, I was moved elsewhere, and there was something of a full-circle moment to be able to come back to a similar project again. While I wish I could be someone who gets to close all their loops, I’ve come to learn that there’s immense value (and satisfaction) in leaving a project in a state where other people can then pick it up and get it over the finish line.

**Implementing your own vision**

I’ve found that the ideal size for an engineering team is four to seven people (although seven is pushing it), with five as the ideal because then everyone has the same closeness to the work, and you always have cover if someone’s off. There’s no point throwing 20 engineers at a problem when what most teams need more than anything is good research and planning—not to mention realistic expectations.

And now that I’m architecting projects, I have a particular philosophy: I’m not only making sure that tasks get done, but also making sure I understand *why* we’re doing what we’re doing. My job is to be the one who asks, “Wait, why are we doing this?” I don’t want to be six months into a project just to find out that we don’t have any data to back up our approach. There’s a lot of trust that I, as a tech lead, need to earn. It doesn’t come just because someone gave me the title of “tech lead”. You have to show that you *are* a tech lead, that you can do it, and you can do it well.

I feel this way partly from previously working in teams with mixed levels of technical understanding, under a manager who would flatly reject proposals from engineers to implement new features. I prefer to be solutions-oriented, and it taught me that you have to give people space to process why something is happening or might need to change, and let them make their own decisions in response. It gives you a middle ground for making compromises as a team.

That brings with it the push and pull of being a partner, not just a stakeholder. Every member of a cross-functional team like ours has differences in technical understanding. It can be fun hearing Product talk about nice-to-haves, but at some point someone like me has to come in and bring everyone back to reality. (No, you can’t just let people type in free text!)

At the same time, since I’m one of the full-stack engineers on the team, I’m aware of how often Engineering can feel like problems are just dumped on its shoulders. It’s been good to flex my muscles both explaining to non-engineers what can and can’t be done, and also explaining to Engineering that we can’t just quit because then nothing’s getting done.

If Design decides something new is in scope, then we have to raise that as an issue. If Product tells us that we’re going to have to work with the Creative Assistant team for example, then we need to talk to them and find out if it’s actually possible. I refuse to leave any conversation feeling like I didn’t do what I needed to do—but it’s not about throwing anyone else under a bus, either. I have to be able to say I put in my best efforts. There’s satisfaction in knowing how much influence that can have on a final product. Having space to have these kinds of conversations is vital.

In a sense, it’s an advantage that Product relies so heavily on Engineering, because it feels like it gives me more of a level playing field when I make my case for changing course. But, sometimes you just have to relent. That’s how you want to do it? It’s not going to end well, but OK!

**Doing by delegating**

There was a point where I realized that I needed to stop pulling tickets myself. The only tickets coming in were the kind that should have been left for someone in a more junior position, starting out in their career. When I’m pulling tickets and working on an API endpoint, I’m not learning anything new, while also ruining a learning experience for someone else.

For example: we were getting the Google Cloud Platform service up and running, and building a new developer workflow for it. I realized it was more helpful—if I was going to pull tickets—to at least record myself, and share those videos with the rest of my team as a resource. I also like to take the product roadmap and make an accompanying technical roadmap, since product roadmaps can get so overwhelming in a feature sense without being clear about what we need to do, when, and why. With our project we had plans to organize, track, and analyze, so I made my own framework alongside: when to spike, track, develop, implement. It makes that key information digestible, plus I don’t have to be involved in every little detail because everyone’s empowered to make decisions.

But I also always leave my calendar open, so I can meet with every engineer on the team at least once every other week for pulse checks. (Plus, Slack’s always there too.) Some people just need someone to talk to for 30 minutes, and I’m happy to be that person.

I hate to see junior engineers starting to burn out. There’s pressure in your early career to keep pushing things out there, and I used to be like that too. (I even had my own existential crisis last year, where I’d ask myself, “If I’m not pushing out code, then what am I even doing here? Should I be an IC? Should I even still be a coder at all?”) But I tell them that life happens, and we’ll be OK if we slow down. Empathy is the key thing—you have to be able to give other people grace.

Something I tell junior engineers is to ask as many questions as they can, no matter how stupid, because at some point in their career they’re going to have to be the person with the answers. Realizing recently that *I’m* that person now was a huge moment, as was hearing from someone that it was my leadership that was helping my team be productive and happy. It’s allowed me to quiet that noise in the back of my head—I can stop focusing so much on the PRs and give people the space and tools they need to get their work done instead.

After all, there are only so many hours in a day, and I can’t be pulling every ticket, helping my team do their work, \_and \_advocating for engineers, all at the same time. And reorienting myself to think in terms of my team instead of just myself applies to my career goals just as much as it does to my productivity and happiness.

**Developing yourself**

If you’re a good tech lead, you can expect to be handed more responsibility—but there’s a balance between wanting to do more, and recognizing that it can take time to reap the benefits of that choice. Manage your expectations, but also set boundaries.

One thing that’s surprised me is realizing that tech leads don’t actually get a lot of time to talk to other tech leads. We have lots of meetings—across Mailchimp and within smaller departments—but there’s very rarely time to just ask someone, “Hey, what’s it like?” That makes it difficult to assess my version of being a tech lead. That said, getting group feedback on tech specs is perhaps my favorite part of the job, when everyone’s reading your work, commenting and clarifying. It’s the closest thing we have to peer review.

I’d argue that there should be an extra title after “tech lead” if someone’s been one for less than a year, because everyone has a different career beforehand and it directly influences what kind of mindset you bring to the role. I know the way I’ve been able to operate, and I’d like to do more of it, but one person’s version of being a tech lead won’t always mesh with another organization’s existing direction, in the same way that it doesn’t mesh when a React developer joins an Angular shop. And if someone asks me what’s next, I don’t know—but I do know that I wouldn’t enjoy this as much if I was just coding, and not able to also strategize and architect.

There’s more to being a tech lead than can be quantified in just the code you put out. You have to prove that your leadership has impact—that you’re closing your loops. Otherwise, what’s the point?

## Breaking down the remote collaboration conundrum

Image Credit: Mar Hernández

*Reorganizing any team can be difficult—but that’s especially true during a pandemic. When Mehdi Vasigh and John “Hutch” Hutcheson found themselves working together on the same engineering team during the summer of 2020, they were struggling with many of their usual processes. They were siloed away in their homes, and so too was the information they needed to be sharing to effectively collaborate. The solution? Developing their own form of remote pair programming.*

*Pair programming isn’t that common at Mailchimp, but left to their own devices, they organically developed a style that worked for everyone on their team. They eventually codified their “rules” in a kind of constitutional document that still guides them. It hasn’t been totally smooth sailing: they had to work with each other’s personalities, handle differences of opinion regarding tools, and the Zoom fatigue is still real. But even if they eventually return to the office, there will still be those who prefer home working—and this experiment has taught them some important lessons about the future of hybrid working.*

*Here, Mehdi and Hutch talk through the journey they took from newly acquainted coworkers to full-fledged collaborators over the course of the last 18 months:*

**Hutch:** Like a lot of office workers, Mailchimp employees started working from home in March 2020, and it took a while for our team to figure out how to work as a unit. I was in the Recommendations squad—a cross-functional team of product managers, product designers, and engineers—working closely with data science on putting smart recommendations in our app. We’d had remote coworkers before, but the majority of people worked in our Ponce City Market offices \[in Atlanta] so it was fairly easy to get up from your desk and bother someone. It became a big issue pretty quickly when we couldn’t just do that any more.

\*\*Mehdi: \*\*And I joined Mailchimp a few months after the pandemic had started, in June 2020. The depth of my remote-working experience was the last three months at my previous company.

**Hutch:** The hardest thing at first was remaining focused—my fiancée and I were living in a studio apartment—but then, secondly, not being able to have ad hoc conversations became a big issue. We ended up with meetings on our calendar for things that could’ve been Slack messages, like “What color should this button be?”

Another engineer and I were having problems combining our work. I would build something, he would build something, we’d try to reconcile them, and then we’d realize we built very disparate things. We were in our own little worlds—me going off and doing the backend work, him doing the frontend work—and it became very clear that what we were doing wasn’t sustainable. Mehdi, I think you were around for the last half of adding Smart Recommendations in the checklist, right?

**Mehdi:** Yeah.

**Hutch:** I built a filtering layer for these recommendations that were coming into Mailchimp, and it would serve the data structure in a particular manner. We’d discussed it on Slack, but when it came time to reconcile it with the UI another engineer had built, we looked at it and were like, “...Ohhh no, this data structure does not fit into this UI, *at all.*” We had to rebuild it all. In the office, if I wanted to ask another engineer something, I could just be like “Hey, can I watch you do this?” versus having to set up a whole meeting. With this information siloing issue, it became very clear that what we were doing wasn’t tenable or sustainable.

### Reach out and touch base

**Hutch:** The way pairing came about was kind of a slow burn.

**Mehdi:** Yeah, from last June up until early this year. It’s very common for engineering teams to parallelize work. Backend and frontend are separate parts, but they have this contract, and sometimes it requires a lot of upfront work to figure out what that contract looks like. But in a small team—like ours has been in my time here—sometimes it’s more productive to be fluid about it rather than try to foresee everything.

\*\*Hutch: \*\*We drafted our eventual [“Ways of Working” doc](https://docs.google.com/document/d/e/2PACX-1vSqxQHcyYK17XiZM1YMcAcH5bRbP5Fz2SuoDtsfWwNC-j3DYlx_Az0g9D0sWtF8FDRpZS3g0WtXrwpd/pub)—or at least the original version—in February 2021.

\*\*Mehdi: \*\*Right. We had an engineer who’s not with Mailchimp anymore—Daniel Swain—and he was a huge part of putting this document together. Making the experience equitable was a big deal for us, especially when there were more than just two people on a call. Defining roles, ensuring that the person writing the code feels comfortable driving, and defining a navigator who’s there to bounce ideas off or direct, so that it doesn’t feel like you’re on a call with a crowd of people telling you how to write your code.

\*\*Hutch: \*\*Us—me, Mehdi, and Daniel—having good personal relationships definitely assisted in having empathy as a group whenever one person felt like something wasn’t working.

\*\*Mehdi: \*\*The process went through multiple passes. We’d do it one way without much planning, figure out what was wrong, go back, talk about it and improve it, and then do it again, iterating it over and over.

\*\*Hutch: \*\*Personal dynamics, professional dynamics, people’s titles can all change—and as those things change, you need to be adjusting your processes, because they’re not going to work forever.

\*\*Mehdi: \*\*And because it was new to me as well, it taught me a lot about myself, and what I could do to build trust with my teammates—everything from making sure that I’m continuously soliciting feedback, to, if I’m navigator, making sure that people are writing their own code and I’m not dictating. Those are all still things that I work on, but it gives you a different kind of space for self-evaluation compared to pulling tickets and pushing commits.

**Hutch:** I mean, there were some selfish motivators—like I didn’t want to be writing PHP all day, and it was convenient to be able to say, “Hey, I want a pair on this,” just so I’m not having to write PHP.

There was a point last year when we really started to flesh out how pairings were going to happen, and I got the opportunity to write code in front of Mehdi and a couple of other engineers, and I remember I started to write a class component, and you and the other engineers were kind of like, “You know you can just write this as a function?” And I was like, “...oh, I had no idea that was a new thing in React!” So, just from a personal development standpoint, it’s also been beneficial to be in front of engineers who are more in the loop.

\*\*Mehdi: \*\*A hugely important part of this was that personality mesh. We have a new team now, and it’s really important to keep evaluating things to make sure they’re still working the way you want.

### We can remote, we have the technology

\*\*Hutch: \*\*I mean, I was the holdout when it came to the tools we were using. I was using IntelliJ, and everyone else was using Visual Studio Code, and I didn’t want to switch…

\*\*Mehdi: \*\*Yeah, a lot of those early sessions were just us ragging on Hutch for IntelliJ being super slow.

**Hutch:** ...but I did switch. And it turned out everyone else was right.

**Mehdi:** That was fun, I enjoyed it very much. Our tools definitely evolved. Late last year we started experimenting with an extension to Visual Studio Code called VS Live Share. It’s really cool, like Google Docs for coding. While it’s definitely got its own quirks and bugs, it’s been a huge improvement because it gives you the option of following someone around and assisting when something comes up. If Hutch is driving and I’m navigating, I can say, “Hey, I’m going to go and write your index file for you,” and do that in the background while he’s working on the meat and potatoes. That’s been huge for us.

Diagramming has also helped us a lot. A lot of the time when we’re starting a new project, we don’t start it in code—we either start it in a Google doc, or Miro.

\*\*Hutch: \*\*Yeah. We’ve developed diagrams in Miro—like merge-tag taxonomies, for example—that we did for our squad, but which have ended up being more widely used by other teams that have attached their work to ours.

**Mehdi:** A lot of the time we sketch ideas out in Miro, drawing boxes and data relationships so we know how we’re architecting something, and how the data is going to flow through a system—“OK, what’s the UI gonna look like for this? Where’s the data going to come from?” Doing that early is a big deal because engineers might have different mental models, subtly or substantially, of how they would architect something, and getting to a consensus is a big step towards not wasting or duplicating effort. Miro’s been a really good tool for consensus  building, not just diagramming.

\*\*Hutch: \*\*A big part of what we’ve been doing that I would like to preserve post-pandemic is to have an artifact that comes out of everything we do. Like with diagramming, traditionally, as Mehdi said, that would be done on a whiteboard at an office, but that’s very temporary. Now we have that document or diagram indefinitely, and we can share it. I’d like to see that stick around, because six months from now when I say, “...what was I thinking?!” it’s much easier to check.

**Mehdi:** At my last company there was an example of a well-intentioned but failed attempt to include remote folks. In Texas, our main team of engineers and data scientists would write ideas on a whiteboard, and it became this completely unwieldy, dense amount of information—and then someone would take a picture and drop it into Slack for folks in California to solve like it was the world’s most complicated puzzle.

It would be great to have a way of organically creating Miro diagrams in an office together where folks can watch live, but it also produces something that has a version history, and it’s cleaner and easier to read than the traditional whiteboard. I wonder if it’s a thing where either the tech isn’t there yet, or maybe it’s the skills. I feel like I’m a little bit better at diagramming now than I was before, but I have experience with those kinds of devices too, and…they haven’t been good.

\*\*Hutch: \*\*We did try Google Jamboards at the old office, but they typically didn’t work out. They ended up being more like that whiteboard experience where everything just seemed jumbled, and thoughts were conflated, and…yeah, not legible.

### Dangerous driving

**Mehdi:** Early on, the way we were pairing was basically just the driver sharing their screen, and if you’ve watched a programmer—like, what they do with their computer—it can be very frantic. For me, I’m switching windows, moving back and forth, all the time. My font size is super small, my theme will be really dark and low contrast, and I realized that I was making people nauseous. Things are in your head, but other folks on the call have to read the code that you’re writing and parse it, so there’s an additional processing overhead.

That was an adjustment I had to make early on. Switching to a higher-contrast editor theme and zooming in to around 150 percent made my code more readable. I also stop frequently to explain my code and to check in with everyone before diving back in. It gives navigators and spectators more clarity, and space and time for thinking of suggestions. Speaking is just as important as typing for communicating your intent clearly.

\*\*Hutch: \*\*And I don’t particularly like being on Zoom calls. It’s not that I \_dis\_like it—it’s just very draining. You do have to be more conscious of meeting fatigue. It’s not something that’s going to go away just because it’s a pairing session, but at least your brain doesn’t actually have to be working as intensely all the time. If I run into a code problem, Mehdi and \[other team members] Erik \[Lopez Ramirez] and Sowmya \[Andalam] are all there to help. It’s a trade-off: there’s going to be meeting fatigue, but there’s also the collective brainpower of your peers behind you.

\*\*Mehdi: **When it comes to the isolation aspect of remote work, I always feel more productive working heads-down. But the communication? That’s been the hardest, especially because when I started at Mailchimp it was really hard to build trust with my team and other folks around me.** \*\*We’ve had some cool department-wide events for remote workers, like when a comedian performed for the engineering and customer teams. We’ve had team-building exercises over Zoom as well, which is fine, even if they don’t really capture the same energy as in real life. The best social interactions were outdoor park hangouts with coworkers, mostly because I feel like the best way to relieve people who have been on Zoom all day is just to give them some time off it.

I think the future of hybrid working presents the challenge of how you combine the in-person and the remote experience so that they’re equitable. How do you create an experience that’s equally accessible to the person in the office and the person in their home? I think that’s a whole new ball game. People need to be very much aware of the folks who are remote, and be able to do everything remote-first, even if they’re in the same room.

### Trust us…we’re engineers

\*\*Mehdi: \*\*One thing that’s maybe a worry for some folks—although it hasn’t been my experience at Mailchimp at all—might be the fear that your manager is going to perceive pairing as two engineers halving their productivity. What’s important is communication: why we do it, and what the expectation is for what we should be able to produce in a given amount of time so that other teams can plan around it.

\*\*Hutch: \*\*Providing a description of the benefits up front is important to that. As an example, when we were writing that “Ways of Working” document, one benefit we pointed out—and it’s still very true today—is that there’s a lot more shared knowledge. If somebody wants an estimate of a timescale, for example, with just Mehdi in the room, they can do that because he has all the information he needs to make an educated statement. I’ve worked on teams that struggled with that, and needed to have every engineer in the room. But now we don’t, which is nice.

\*\*Mehdi: \*\*Pairing has definitely given us more confidence in that way. And importantly, people receive praise equally, too. We get praised as a team, and we give praise as a team. I enjoy working in environments where people on a team aren’t competing with one another, and I think pairing helped facilitate that as well.

\*\*Hutch: \*\*I concur with pretty much all of that. I don’t feel like I have to represent backend engineering at every meeting—I trust the other engineers on my squad. On previous teams I’ve been worried that if I go on vacation I’m gonna come back to mountains of work, and that’s very much not been the case for us.

\*\*Mehdi: \*\*So pairing is definitely not for every team, and there’s probably a universe in which it wasn’t the right thing for us—if we were working with different people, or on different things. I think the overt pressure of “You have to pair on everything” is kind of mortifying to me, honestly. I don’t want to be in that kind of environment either. It’s more nuanced. Pair programming isn’t supposed to be dogmatic, so don’t go and tell your organization, “We need to pair on *everything*, it’s going to make everything better!” because it’s not true.

\*\*Hutch: \*\*Yeah, you do need to be constantly evaluating whether this is right for you, and right for everybody. If an engineer chooses not to participate in a pairing—whether because it doesn’t suit their working style, or they just have Zoom fatigue—that should be OK. On our team, when we set up a pairing meeting, it is not a hard-and-fast rule that you need to be there. It *should* be that everyone wants to be there, but that’s not a requirement.

\*\*Mehdi: \*\*It really comes down to the needs of your team around you, and then also evaluating whether it’s getting you the solution that you want. Every team is different, and there are many ways to create a collaborative environment where information is freely shared. For your team, it may be pair programming, producing and sharing artifacts, thorough documentation, or something else entirely. It’s worth investing the time to craft and refine your team processes.

## Changing the tires on a moving bus

Image Credit: James Daw

As an engineer at Mailchimp, one of our responsibilities is to do a tour of duty in the on-call rotation roughly once per quarter. This is a full 24-hours-a-day, 7-day rotation where anything falling apart outside of normal business hours gets escalated to you and your team. While you’re keeping track of everything for that week, someone will occasionally hit you up for a favor and ask you to check in on the progress of something. So we begin our story there.

“Hey Nate,” a support agent said to me. “Would you mind keeping an eye on these account exports over the next few days since you’re on call?”

I started working at Mailchimp in 2010, back when we were a much smaller company—my employee number is in the double digits. Even though I’d gone to school to be a history teacher and my résumé was full of jobs where I had my name stitched into my shirt, my first salaried job was here at Mailchimp as a support agent. I’ve changed job titles half a dozen times since then, but Support will always be home to me. Obviously when someone asks a favor I try to be true to my word. But when it’s my home team, I go out of my way to make sure I follow up on what I’ve committed to.

So of course I said, “Yo! Absolutely.”

At this point in my tenure, the account export feature was more than a decade old. That code had existed since *I* was in Support, so I’d naively assumed things were working as expected all the time. The truth was, as Mailchimp grew, account exports had become a growing pain in the neck for our Support teams: users would begin an export, but if that export was unusually large, the export job could crash and disappear, and the user would be left in the lurch.

As a result, our Support teams devised a way to estimate when a user’s account was going to become too unwieldy to export without substantial intervention. They based it on a number of factors, but essentially, the larger the account, the less likely it was to complete successfully on its own. The export jobs Support had asked me to watch were deemed “at risk” by these criteria, so they figured an extra set of eyes wouldn’t hurt.

For the first few days, I stared at some graphs that were meticulously logging every operation being conducted during this user’s account export—real riveting stuff. To the surprise of no one in Support, the exports they asked me to watch failed miserably about halfway through. I noticed that even though the exports failed at different points of execution, they’d both failed within a few minutes of each other. And while nothing stuck out about the job itself, I did notice that we had restarted our job-management system right around the time the exports failed.

Here’s what’s supposed to happen: a user clicks the button inside the application that says—straightforwardly enough—“Export data,” and then we email them when it’s ready. On the back end of things, we are furiously pulling data from our databases and compiling it into an array of CSV, HTML, and image files of all types, asynchronously. Each account is slightly different due to variations of sending habits and size, so the fact that these two distinct account exports had failed so close together meant \_something \_to me.

With the timing clue, I was able to track the problem down to a job-management-system restart: as part of that restart, we were deleting jobs that had been running for more than 24 hours. While I was sure we’d originally started doing that for a reason, to me, it was only creating problems, pinching our export jobs with no remediation, and leaving no evidence as to why they would have failed.

Since I was on call, had a pretty good mental model of the situation, and had a few days of time where I could prioritize this project, I decided that I alone could fix this problem. Dear reader, if you ever find yourself thinking this way, let this be a tale of warning to you: this “fix it in a few days” adventure turned into a full year of tinkering, learning, and breaking things in new ways before finally settling on a solution that was to our liking.

### A first attempt at fixing the problem

Feeling very smart and full of the optimistic vigor you get when you finally uncover a perplexing bug, I decided to rewrite our export process right then and there. In hindsight, this was an incredibly optimistic and naive approach. Today I—an older, wiser, more weathered engineer—would pump the brakes on such wild timeboxing.

Our exports were getting killed more often than other jobs because of the sheer length of time they took to run. Some of our largest users’ exports could take anywhere from several days to several weeks to complete. These were single jobs that were just building and crunching data, uninterrupted for that entire time; there was no way to pause or resume, so failure meant starting all over again. In my mind, the fix seemed simple: make jobs responsible for one piece at a time, make them shorter-lived, and make more of them.

Originally, it was just me on this project. But as I got further down the rabbit hole, I began to vocalize my intentions, and thankfully, another courageous fool decided to join me. Enter my teammate: Bob. Once Bob and I got to coding and then testing in our staging environment, things were working as expected. Smaller jobs? Check. Building a zip file in our testing environment? Check. Gratuitous, self-congratulatory back-pats? Buddy, you’ve never seen so many.

But this story wouldn’t be fun without a tale of self-sacrifice. When our customer support team gave me the go-ahead to run our newly forged account export process on a test user, we only got half the data back. Perplexed, I ran it again, thinking there must be a glitch in the Matrix (you know how computers are). When I got back a slightly different but still half-complete data set, I knew something was up.

I logged on to one of the production servers this user resided on and found the other half of my missing data. Our first major roadblock: I had failed to take into account the sharded nature of our production environment. Mailchimp runs on multiple hosts, and for a distributed job system, we have to take the physical occupancy of the data into account.

To give a concrete example, imagine you and a friend are writing a story where you alternate every other sentence, but each on your own sheet of paper. At the end of the story, the teacher calls on you to read the story aloud, but you only have access to the sentences you wrote on *your* paper.

That’s what was happening with user data: when the job runner called the job to zip things up, only the data on the selected server was available. Most of the time, this meant a fragmented export result. Our customers expect to receive *all* of their data that they export, so we were back to the drawing board.

To rectify this behavior, we needed a central repository where we could store the data and have it available, agnostic to which server needed it. We eventually decided to cast our lot with Google Cloud Storage (GCS). With a little elbow grease, our jobs were uploading data into GCS and downloading it back on the server that needed it to complete the export.

Imagine you and your friend are writing that collaborative story, but now you’re using the same sheet of paper—you can see the whole thing instead of just your own lines. Exports were now running in a much smoother and more reliable way, and we were giving users all of the data they requested. This was “mission accomplished!”

### Mission not accomplished

Fast forward a few weeks. Another set of engineers are on call. In the middle of the night, they get a page for one of our servers running out of available disk space. They traced it back to a particular user and…our export job.

If left untreated, this job threatened to exhaust all available disk space and cause some very nasty downstream effects for everyone else. In order to save the rest of the users on the shard, they killed the export and reached out to the user to let them know. In the end, a decade’s worth of open-and-click data, CSVs full of list contacts, and all of the files in the user’s image gallery ended up being larger than the physical disk space available on our production host.

So what had gone wrong? Well, in a way, nothing “went wrong”—things were working exactly the way we designed. (And if you happen to be my manager reading this, you might say they were working \_too \_well because we did *too* good of a job!)

The problem was that we never expected the floodgates to open this much. Users who had previously never been able to create a successful export were suddenly having their entire accounts’ worth of data dump out onto our servers. Because these were generally one-off situations, we triaged them as best we could, even breaking up our exports into sections (thanks to our new job structures) to give the users the data they needed without melting things.

It wasn’t until our Customer Success Team reached out to us that a large media organization had begun to export a decade’s worth of Mailchimp data that my coworker Bob and I really realized how deep we were in this. In the fog of war, we decided we would insert a code shim for this user that would prevent us from downloading that data and almost assuredly wrecking all available disk space. From there, we’d figure out a way to get our data out of GCS and give the user a hand-crafted/bespoke/hacked-together version of their data.

### Okay, but like…how?

The idea of giving this media conglomerate their data in a nice, easy-to-parse GCS bucket seemed great, but the reality was that we had to assemble it. Google doesn’t allow you to grant access to a subdirectory of data; it’s more of an all-or-nothing-type approach.

So how were we going to do this? Bob had an idea to use Go as a (pardon the pun) go-between for our batch server infrastructure and our GCS bucket. We’d reach out to the bucket, compress the files into a zip format, and then move our newly formed zip file back into our GCS bucket. Then we could serve them their data in a secure, signed, and obfuscated URL that was hidden behind our authentication. That checked a lot of our boxes for getting them their data in a way that made sense.

With our immediate fire out of the way, we began to pursue this lead to its (il)logical end: What if we did this for every export for every user? There were some obvious wins right out of the gate: we could save money by not using another storage host for the old export files, we could have all of our user data in one secure and predictable place, and since we would no longer have to write to the disk, we wouldn’t knock over our servers when processing a gigantic amount of data.

### Getting the gang together for one last heist/refactor

Wiring this up for a special user was one thing, but wiring it up for everyone who uses our account export feature was quite another. We had dependencies, legacy code, and user expectations to consider. What if someone had just completed an export, and it was being stored in our old place, and we tried to link them to our GCS file instead? We either had to be very opinionated or at least medium-smart about how we did this.

With a little trial and error, we were able to settle on a slow-rolling feature-flag situation where, as new exports were queued up, we enabled this new GCS-hosted option instead of our previous host. This meant that any exports that were in-flight wouldn’t have the rug jerked out from under them, nor would previously completed exports.

But this left us the problem of how to eventually migrate away from our previous host completely. Users still had completed exports, and we didn’t feel great about just having those disappear without any sort of warning.

This is where we had to be opinionated but empathetic towards our users. We decided to only allow export downloads for a period of time before they would eventually disappear or need to be restarted. This helped us to “roll off” our previous host slowly. Account exports are a free feature for our users—aside from the investment of time spent waiting—so this didn’t feel as bad as it could have.

The technical details of our implementation were actually pretty plug-and-play for a feature that was almost a decade old. Instead of uploading this file to our old host, we’d queue up our Go binary to act as the interface between our batch infrastructure and GCS. This would zip their file, sign the URL, and save it in our DB for quick recall. With our retention rules in place at the Mailchimp level, this allowed us to implement retention rules on the GCS side as well, eventually rolling off old data that we no longer served up to the user.

### Happily ever after

When I initially waded into this project, I assumed it’d be resolved in a matter of weeks. I guess that’s still technically true—it just took 52 of them. What started out as a passion project with a skeleton crew of myself, one other engineer (shout out again to Bob!), and the grace and air cover from our various project and engineering managers, we were able to turn a neglected, decade-old feature into a more modern and polished option.

In 2011, account exports at Mailchimp were a slow, monolithic process that exported all of your data in one go. A decade later, we allow our users to pick what data they want and the timeframe they want the data from (30, 60, or 90 days), and we have hardened it against our most common failure points. That’s not to say that we don’t still have more work to do, but maybe I’ll save that refactor until 2022.

## Scaling in the age of COVID

Image Credit: James Daw

Think back to March of 2020—I know it was decades ago at this point. It was the early days of the pandemic. Social and business activities had shifted online, but the migration was sudden and abrupt; it would take another decade (read: month) for the world to name the fatigue.

Products that support collaboration and socialization experienced a sharp rise in usage in March of 2020—and so, in turn, did products that support those products. I do scaling and modernization work for one such product: [Mailchimp Transactional](/transactional), which lets customers send one-off emails like password resets, order confirmations, and meeting reminders.

As our clients saw that sharp rise in usage at the onset of the pandemic, so did we. For backend engineers like myself, more usage is simultaneously exciting and scary. We usually design infrastructure to support a particular scale and pattern of usage, and changes to either of those things can impact the performance of critical systems.

But the spikes we saw at the start of the pandemic were unprecedented: both a general increase in traffic *and* a dramatic increase in usage for a small number of users. Like many web applications, Transactional’s database layer is most impacted by load changes, because most operations taken by users result in a need to read or write data. Any system that relies on the main database (so: all of them) might experience a partial outage if heavy traffic causes the database layer to be slow or otherwise unresponsive.

We spent our early pandemic days firefighting outages that were mostly caused by the strain on our database as it worked to keep up with the newly elevated and asymmetric load. We began monitoring the systems assiduously, which gave us advance warning of performance degradation. As the year went on, we moved from a period of unexpected interruptions to a period of constant vigilance. Both of these states are tiring and not desirable, but the latter gave us the control necessary to be proactive.

More than a year later, the highest-volume Transactional users are still sending significantly more emails than our average customers. But while our overall usage continues to be high compared to pre-pandemic levels, we’ve improved the health of our database and given our on-call engineers some relief.

Here’s what happens when a perfectly reasonable scaling scheme hits a sudden change in usage patterns—and how we adapted to it.

### Mailchimp Transactional: When you’re busy, we’re busy

Mailchimp Transactional is a service for sending transactional email. Most Transactional customers are developers who embed calls to our product from within their own applications so they can trigger emails in response to actions taken by their own users—things like password reset requests, login notifications, and more.

Transactional is a product that supports other products. For most software engineers, we’re a third-party service that provides a reliable way to communicate with their own users. As the pandemic drove social, educational, and business activities online, companies that host those activities needed to send correspondingly higher volumes of one-time passwords, account verification messages, and other email-based notifications—and we needed to continue providing the same standard of delivery speed that we’re known for, despite the higher load.

Originally called Mandrill, Transactional was created nine years ago, and until last year, it operated primarily with its original architecture. Software is not eternal, and in 2019, Mailchimp leadership committed to modernizing and maintaining the product. But, of course, they had no way of knowing what the next year had in store.

If 2020 had not been a global trash fire, we probably would have spent more time optimizing some of our core technologies, which could have sped up our platform and allowed us to upgrade some of our third-party dependencies. We still need to do that work, but in 2020, what our users needed most was a stable platform, not a slightly faster one. As a non-product-facing team, we had the freedom to adjust our roadmap based on our users’ immediate needs, and we were able to instantly pivot to addressing our emergent scaling problems.

Scaling work is well within our team’s purview. Our team of application and infrastructure engineers spent the past year upgrading Transactional’s hardware and software dependencies and rearchitecting systems that did not meet our reliability expectations in light of the higher load or that weren’t scaling well. We upgraded our database and Elasticsearch integration and improved the performance of a few features by multiple orders of magnitude.

But we also found ourselves sinking hours into babysitting the scheduled sending processes and manually cleaning up Postgres deletion flags. And, despite our best efforts, the performance hits were occasionally noticeable to our users. At the end of the day, the gold standard for application stability is happy users, so we stalled our original agenda to investigate why our application was struggling to handle our increased scale.

### Built to last...for “ordinary” use

Part of why Transactional was able to operate for so long without major, scale-related failures is that the original developers *did* anticipate increases in user numbers and sending volume—and their strategy worked for eight-plus years, which is basically a century in web infrastructure time.

The original developers envisioned a future in which the number of users would grow, but the sending volume wouldn’t vary too much across the user base. Some users would always be higher-volume than others, but they assumed that future users would be developers at companies of comparable size, using the platform to send comparable volumes of email.

They also knew that if the product grew successfully, the amount of data they’d someday need to store would outgrow the maximum volume of a single server. With that future in mind, they implemented sharding by user ID for most data.

Sharding is a technique used to distribute large data tables into smaller chunks that can be hosted on separate servers. Without sharding, all new rows for a table would get written to the same table on the same server. Eventually, that table would get very large, and disk space on the server would fill up. To allow the table to continue to grow, you could upgrade to a larger server—this is called vertical scaling. Alternately, you could split the table among many servers—this is called horizontal scaling. Each segment of the table is a partition/shard, and sharding is how you decide which partition each row is stored on.

A typical strategy is to hash a characteristic of the row—usually a column value or combination of values—modulo the number of shards. Every row with the same shard key (the value that gets hashed) ends up on the same partition, so shard keys are frequently chosen such that common read operations only need to access one shard. Like many web applications, Transactional frequently needs to fetch data for specific users, so sharding by user ID is a natural choice.

(Although our sharding strategy is fairly typical, our nomenclature is a little nonstandard. We call partitions “logical shards,” and they’re implemented as separate schema in Postgres. These logical shards are distributed evenly across physical database servers, which we call “physical shards.”)

Here’s where the assumption that all users have a similar amount of data comes into play. The sharding algorithm distributes users evenly across the logical shards, so if all users have about the same amount of data, sharding by user ID will spread that data evenly across logical shards. Because each physical shard has the same number of logical shards, the resulting load should be distributed evenly across physical shards.

For years, that assumption was a pretty good one. Although the product was successful and our users sent a lot of emails (thank you!), we didn’t have superusers with outlying amounts of activity.

But the COVID-related increase in online activities drastically changed the distribution of usage across our user base. The most active Transactional customers now had orders of magnitude more data than the average user. The logical shards that housed superusers’ data were getting significantly more reads and writes than the other logical shards, putting extra stress on their physical shards. Physical shards with hot logical shards have higher CPU load and higher memory usage, which makes them less performant.

![In the early days, we had a lot of headroom on each database shard. When the pandemic hit, we started to see some shards carrying disproportionately large amounts of data.](https://images.ctfassets.net/cxsachrr7h0p/1NmxJw8mhxft4nJI2rTDoh/068196ab7eb72a9971db350e10e21c85/Silverio_01.png)

*In the early days, we had a lot of headroom on each database shard. When the pandemic hit, we started to see some shards carrying disproportionately large amounts of data.*

Nobody wants a database to struggle. Basically every operation in a web application relies on the database layer at some point, so our overloaded shards were concerning. But in the end, it was an unfortunate coincidence that pushed our concerning situation over the edge into a source of frequent fire drills.

### What’s the likelihood of a hash collision anyway?

So, we had logical shards that were struggling to keep up with the sending volume of our new superusers. These logical shards were hot because they house large instances of a table related to our search feature, which lets users examine their recent activity and requires one database row per message sent. The Search table is quite large (we send a lot of emails) and because it’s directly tied to user activity, that table shows the most variation in size per shard.

When I plotted table size per logical shard in April of 2020, I found that the mean size of the table was twice the median, indicating that the mean was strongly impacted by large outliers. Because the Search table is generally quite large, the logical shards containing the outlying Search tables also consumed an outlying amount of disk space.

We also store our job queues in a database table in the main database, sharded by the queue identifier. Even before the pandemic, I would have preferred that our job queues were not in our main database. The Job table has a lot of reads and writes, and most jobs initiate reads and writes to other tables, so a busy job queue that causes elevated load on the database can start a feedback loop if it’s stored on the same database it’s straining.

As a still-growing team, we hadn’t had bandwidth to change that yet, and we expected that the system would hold for the time being. Unfortunately for us, once the pandemic hit, the logical shard that contained the largest instance of the Search table happened to be on the same physical shard as the busiest job queue.

That shard became known as a Problem for our on-call engineers. With the elevated pandemic traffic, the Problem shard was continuously operating near a dangerous performance threshold, below which users would notice unusual latency for common operations. Slight variations in traffic could push shard health over the threshold, disrupting users’ normal experiences (and paging us!). We worked hard throughout the summer to minimize customer disruption, quickly responding to pages from various monitoring systems. As anyone who has been on an on-call rotation can attest, keeping a product operational is deeply rewarding and also deeply exhausting—we were highly motivated to address the high database pressure.

Unfortunately, even operations that would normally relieve database pressure also occasionally worked against us—routine cleanup of delete flags, which reclaim disk space taken up by records of deleted rows, struggled to complete on the giant Search table. The extra resources taken by those cleanup operations would push the shard into dangerous territory, and every time we tried to repack Search, the scheduled send queue would back up.

Our team spent most of the spring and summer putting out fires caused by the non-performant database server. Operations involving the Search table in particular became less reliable for users on the impacted physical shard. Event-tracking systems also use the Search data, so the unreliability of the Search table occasionally resulted in cascading delays of open-and-click processing. The busy scheduled send job queue became prone to backups that correlated with database health, and would resolve itself as soon as the physical shard’s CPU dropped back to an acceptable threshold. We are proud of our response times and dedication to our customer base during this period of firefighting, but clearly needed to resolve the core problem.

![We were confused about job queue length fluctuations that had seemingly nothing to do with our actions…until we realized they correlated strongly with changes in database CPU.](https://images.ctfassets.net/cxsachrr7h0p/3PO60rw57RQ1E4wvarE1C6/bd17a13b09c95f55c7b0b7918915f67f/Silverio_02.png)

*We were confused about job queue length fluctuations that had seemingly nothing to do with our actions…until we realized they correlated strongly with changes in database CPU.*

From an outside perspective, these systems (Search and scheduled sends) are not strongly connected, making the correlated symptoms puzzling for users. While we worked to set up comprehensive dashboards and metrics, the behavior was occasionally puzzling to us, too. We’d see a job queue back up and then mysteriously “resolve,” not realizing that the backup and draining of the queue correlated with the beginning and end of a background process that used the Search table.

### Let’s put this fire *away* from the rest of the fire

Turns out, this isn’t a particularly difficult problem to solve. The primary problem is that two hot datasets live on the same server, so we ultimately need to separate the datasets. From a software-architecture perspective, it’s the path that will ultimately give us the most runway to grow. If we had been game for a quick-and-dirty solution, we could accomplish separating the hot datasets by forcing a shard rebalance—in other words, by changing the distribution of logical shards over the physical shards.

However, neither dataset is particularly well-suited to be in the main database. The scaling requirements of the Search dataset are very different from the rest of Transactional’s data, which don’t vary as strongly with user ID. Search data also expires, which requires frequent deletes. In Postgres, the DBMS we use, deletes are implemented by saving the deleted row with a delete flag, and then actually deleting the flagged data later with another process called `VACUUM`. Most other tables in the main database don’t have such high turnover; Search, with its high variance and constant need to be vacuumed, is an oddball.

Job queues have very different read and write patterns from general application data, meaning that pulling them out of the main database would also bolster the database’s health. Unlike Search, however, job queues are a poor fit for relational databases in general. They have high I/O and a short lifespan, and they’re usually generated from some main application code and then picked up by a pool of independent workers, so they’re a good fit for pub/sub systems. From a software architecture perspective, both the Search table and the job queues should be moved to separate datastores. Due to the relatively untouched nature of our codebase, those refactors are quite large and require significant discovery work as well as actual implementation effort.

![Isolation helps prevent load-intensive processes from impacting each other: separating Search data and job queues will ensure that load changes will remain localized.](https://images.ctfassets.net/cxsachrr7h0p/GJ6SGnybQMZjIUpuYVegG/c6121cc95a9bb88e10028b2eef28fb7a/Silverio_03.png)

*Isolation helps prevent load-intensive processes from impacting each other: separating Search data and job queues will ensure that load changes will remain localized.*

Our team was primarily composed of operations engineers, with only a handful of application engineers, which makes infrastructure changes much easier for us to execute than large-scale refactoring of the application. The extreme sensitivity of our core systems over the summer made immediate action necessary, so our operations engineers bumped all of our Postgres servers to upgraded hardware.

That improvement resulted in immediate relief for the hot shard. The busy job queue is now much more resilient, the event-processing queues are much less likely to churn, and Search operations are performing well. We’re still addressing the primary problem, but the fast response of our operations engineers has bought us time to plan and implement the large-scale refactor.

### Is the moral of the story to just throw hardware at the problem?

I would love to be able to say that the migration is already complete. Horizontal scaling (spreading across more hosts) is a longer-term solution than vertical scaling (bumping host size), so until Search is in a separate database, we still have room to optimize.

But as of the time of this writing, the migration is well underway. Our work thus far, focused on restructuring the ORM so it can tie in to another datastore for Search, has primarily resulted in rebuilding lost institutional knowledge of Transactional internals. While it’s been frustrating at times to spend a lot of time laying groundwork in a hard year, we’re taking solace in the fact that the groundwork will help us accomplish the eventual migration with confidence, and both Search and scheduled sends are performing, with minimal intervention from us.

So yeah, we threw hardware at the problem, and it’s helped a lot. But the true moral of the story is, in my opinion, nontechnical. We’ve shown that it’s possible to improve Transactional—and we’ve shown that delivering a series of incremental improvements can reduce toil and establish breathing room for larger-scale refactors.

At the end of the day, all software becomes legacy as soon as it is deployed. Our users’ needs may have changed suddenly, but they would have changed someday, anyway, and it’s our job to listen hard and change our system fast.

## Computers are the easy part

Image Credit: Mar Hernández

In the world of aircraft safety, a Controlled Flight Into Terrain (CFIT) is an accident where an aircraft that has no mechanical failures, and is fully under the control of its pilots, is unintentionally piloted into the ground. As a concept, CFIT has long been studied to try and understand the human factors involved in failure. The [FAA reports](https://www.faa.gov/news/safety_briefing/2018/media/SE_Topic_18-11.pdf) that contrary to what one might expect, the majority of CFIT accidents occur in broad daylight and with good visual conditions—so how, then, is it possible that highly trained and skilled pilots could accidentally fly a plane into the side of a mountain?

It’s tempting to ascribe such a tragedy to human error. But the work of researchers like Sidney Dekker challenges this view, instead framing human error as a symptom of larger issues within a system. “Underneath every simple, obvious story about ‘human error,’” Dekker writes in\_ \_[*The Field Guide to Human Error*](https://books.google.com/books?id=NTP5iLn4XHkC\&q=underneath#v=snippet\&q=underneath\&f=false), “there is a deeper, more complex story about the organization.”

These systemic faults—often cultural in nature, rather than purely technical—are how a group of highly skilled individuals reacting rationally to an incident could nonetheless end up taking the wrong course of action, or ignore the warning signs right in front of them.

We don’t pilot aircraft at Mailchimp, but millions of small businesses do rely on our marketing platform to keep their businesses running. When things do go wrong—and as we work to fix them and analyze what happened—we run up against similar questions about technical versus systemic failures.

We recently experienced an internal outage that lasted for multiple days. While it fortunately didn’t impact any customers, it still puzzled us, and prompted a lot of introspection. Human factors, the weight of history, and the difficulties of coordination caused this issue to stretch out far longer than it could have, but it taught us something about our systems—and ourselves—in the process.

### The investigation

It was the late afternoon on a day where many of our on-call engineers were already tired from dealing with other issues that we started receiving alerts in the form of what we call “locked unstarted jobs”—essentially, individual units of work being claimed for execution but never run. This is a particular failure mode that is well known to us, and while uncommon, has a generally well-understood cause: some kind of unrecoverable failure during the execution of a task. Our on-call engineers began triaging the issue, first trying to identify whether any code had been shipped at the time the incident began that could have caused the problem.

Incident response at Mailchimp is transparent to all of our employees—when we become aware that something’s wrong, we spin up a “war room” channel in Slack that the whole company can observe. The responding engineers, based on our prior experience dealing with this type of failure, first suspected that we’d shipped a change that was introducing errors into the job runner and began mapping the start of the issue against changes that were deployed at the time. However, the only change that had landed in production as the incident began was a small change to a logging statement, which couldn’t possibly have caused this type of failure.

Our internal job runner—which executes a huge variety of long-running tasks asynchronously—is a long-serving part of our infrastructure, developed early in Mailchimp’s history. It runs huge numbers of tasks daily without very many issues—which, on the surface, is exactly what the operator of a software system wants. But this also means that collectively, we don’t often build the expertise to debug novel failures, compared to the battle scars that engineers develop on systems that fail more regularly. The job system has a handful of well-understood failure modes and over the years, we’ve developed a collection of automations and runbooks that make rectifying these issues a routine and low-risk affair.

When a war-room incident starts up, on-call engineers from various disciplines gather together to take a look. Having been intimately familiar with the various quirks of the job system over the years, we had a reasonably solid mental model that when a failure of this nature occurs, it’s generally an issue with a particular class of job being run or some kind of hardware issue.

Through our collective efforts, we were able to quickly rule out hardware failures on both the servers running the job system and the database servers that support it. Further investigation didn’t really turn up any job class in particular that might be causing issues. But by late in the evening, we still hadn’t had any breakthroughs, so we put a temporary fix in place to get us through the night and waited for fresh eyes in the morning.

### Fresh perspectives

By the second day, the duration of the incident had attracted a bunch of new responders who hoped to pitch in with resolution. Based on what we’d seen so far, we had ruled out any obvious hardware failures or obviously broken code, so the new responders began investigating whether there were any patterns of user behavior that might be creating problems.

As we began to dig into user traffic patterns, we noticed a number of integrations that were generating huge numbers of a particular job type. We attempted a change to apply more back pressure for this job class to see if it would mitigate the issue, but that didn’t really help.

The numbers of locked and unstarted jobs continued to climb, and we realized that this was a failure mode that didn’t really line up with our mental model of how the job system breaks down. As an organization, we have a long memory of the way that the job runner can break, what causes it to break in those ways, and the best way to recover from such a failure. This institutional memory is a cultural and historical force, shaping the way we view problems and their solutions.

But we were now facing a potentially brand-new type of issue that we hadn’t seen in a decade-plus of supporting the job system—it was time to start looking for a novel root cause. We began adding more instrumentation to the job system in an effort to find any clues that we’d overlooked in the first day of the investigation, including some more diagnostic logging to help trace any unusual failures during the execution of specific jobs.

### A breakthrough

With this new instrumentation in place, we noticed something incredibly strange. The logging that had been added included the job class that was being executed, and some jobs were reporting that they were two different types of job at the same time—which should have been impossible.

Since the whole company had visibility into our progress on the incident, a couple of engineers who had been observing realized that they’d seen this exact kind of issue some years before. Our log processing pipeline does a bit of normalization to ensure that logs are formatted consistently; a quirk of this processing code meant that trying to log a PHP object that is [Iterable](https://www.php.net/manual/en/class.iterator.php) would result in that object’s iterator methods being invoked (for example, to normalize the log format of an Array).

Normally, this is an innocuous behavior—but in our case, the harmless logging change that had shipped at the start of the incident was attempting to log PHP exception objects. Since they were occurring during job execution, these exceptions held a stacktrace that included the method the job runner uses to claim jobs for execution (“locking”)—meaning that each time one of these exceptions made it into the logs, the logging pipeline itself was invoking the job runner’s methods and locking jobs that would never be actually run!

Having identified the cause, we quickly reverted the not-so-harmless logging change, and our systems very quickly returned to normal.

### Overlooking the obvious

From the outside, this incident may seem totally absurd. The code change that immediately preceded the problem was, in fact, the culprit. Should have been obvious, right? The visual conditions were clear, and yet we still managed to ignore what was right in front of us.

As we breathed a collective sigh of relief, we also had to ask ourselves how it took us so long to figure this out: a large group of very talented people acting completely rationally had managed to overlook a pretty simple cause for almost two days.

We rely on heuristics and collections of mental models to work effectively with complex systems whose details simply can’t be kept in our heads. On top of this, a software organization will tend to develop histories and lore—incidents and failures that have been seen in the past and could likely be the cause of what we’re seeing now. This works perfectly fine as long as problems tend to match similar patterns, but for the small percentage of truly novel issues, an organization can find itself struggling to adapt.

The net effect of all of this is to put folks into a “frame”—a particular way of perceiving the reality we’re inhabiting. But once you’re in a frame, it’s exceedingly difficult to move out of it, especially during a crisis. When debugging an issue, humans will naturally (and often unconsciously) [fit the evidence they see into their frame](https://www.researchgate.net/publication/220579480_Problem_detection). That’s certainly what happened here.

Given the large amount of cultural knowledge about how our job runner works and how it fails, we’d been primed to assume that the issue was part of a set of known failure modes. And since logging changes are so often completely safe, we disregarded the fact that there was only one change to our systems that had gone out before the incident started—had that change been something more complicated, we might have considered it a smoking gun much sooner.

Even our fresh eyes—new responders who joined the incident mid-investigation—tended to avoid reopening threads that were already considered closed. More than once, a new responder asked if we’d considered any changes that shipped out at the start of the incident that could cause this, but hearing that it was “just a logging change,” they also moved on to other avenues of investigation.

We all collectively overlooked the fact that complicated systems don’t always fail in complicated ways. Having exhausted most of our culturally familiar failure modes for the job runner, we weren’t looking for the simplest solution; we assumed that a piece of our infrastructure as old as the job runner must have started exhibiting a novel and unknown problem.

Doing our incident response in full view of the entire engineering team was critical here, since it enabled us to attract the attention of people who were familiar with this obscure type of failure—but this also further highlighted the trap we’d fallen into. With the benefit of hindsight, the responders could have started their investigation from first principles, or tried reverting the logging change as the initial and simplest explanation. Instead, we needed to be bailed out by folks who happened to have seen this exact kind of failure in the past—which is not a resolution that an organization can count on all the time.

This was an incredibly valuable reminder for us: the weight of history and culture within a software organization is a powerful force for priming individuals to think in particular ways and can result in difficulty adapting to novel problems. Approaching incidents like this from first principles and starting with the simplest explanations—no matter how likely they seem—can help us overcome these kinds of mental traps and make our response to incidents much more flexible and resilient.

## Empowering developers to empower the underdog

Image Credit: Mar Hernández

I was bleary-eyed, sipping coffee on yet another pandemic morning in the autumn of 2020, when I received a message that grabbed my attention. The subject line read: “*Help Mailchimp empower the underdog!*”

Mailchimp, I soon learned, was ramping up their investment in developers, and they wanted to bring that to the next level by hiring a Developer Advocate. Admittedly, I wasn’t looking for a new role—and maybe more importantly, I’m not actually a Developer Advocate. But I was curious, and empowering the underdog is my jam. I decided to talk to them anyway.

In my experience, a Developer Advocate represents the external developer community internally, and the internal developer experience externally. Kind of a mental tongue-twister! Developer Advocates need the right tools to succeed: an internal API strategy tightly linked to Product, for example, or quantitative and qualitative metrics on how developers are working with APIs and other products. How happy are developers with API documentation, or the ability to troubleshoot integrations while building, or getting inspired on what to build in the first place?

My first conversations with Mailchimp made it clear that developers were an essential part of their overall growth as a company. They’d brought on a stellar agency to relaunch [the developer site you see today](/), and they’d invested in a complete overhaul of their documentation.

But there were still some major gaps in their strategy. How did they plan to expand their APIs as part of the product roadmap? How were they going to approach their growing community of app builders? And how could they elevate their incredible engineers and showcase how interesting and fun it is to work at a company that processes one billion emails for their customers every day?

Throughout my career, I’ve been lucky to find myself at some pretty amazing companies addressing these very same questions. With a background in B2B marketing, I pivoted to [Developer Community and Developer Relations](https://www.marythengvall.com/blog/2019/5/22/what-is-developer-relations-and-why-should-you-care) a decade ago, when I took on a “wearing all the hats except directly engineering the API” role at Context IO. I would go on to build foundational developer ecosystems for companies like Intel and Shopify. After a stint in consulting, my last role was at HubSpot, so clearly this whole small business platform jazz is appealing to me.

There’s no company that cares as much about small businesses as Mailchimp. But they also didn’t realize that they weren’t ready for a Developer Advocate. So I somehow chatted my way into an offer for a new role: Senior Manager of Development Community, with a focus on building out the structures a Developer Advocate would need to successfully listen to the developer community and create resources based on their direct feedback. Helping Mailchimp build the foundation that will *empower developers to empower the underdog*—our Developer Community mission—was too good an opportunity to turn down.

### Listening to developers first

When many organizations start to build developer programs, they only look outwards, at the developers using their products. Shockingly few start inward—with their very own software engineers—but there’s actually no better place to start.

At Mailchimp, our engineering mission is to “give marketers production-ready software designed to help them grow” and engineering success happens through “togetherness, momentum, and pragmatism.” Our engineers have built some incredible developer products over the years. Beyond our core [Marketing API](/marketing), Mailchimp launched Mandrill—now [Mailchimp Transactional](/transactional)—nearly a decade ago; it currently powers email operations for thousands of developers and organizations. And in 2020, Mailchimp acquired Reaction Commerce—now [Mailchimp Open Commerce](https://mailchimp.com/developer/open-commerce)—an open-source-licensed commerce platform that allows developers to customize its API and build commerce stacks for retailers of any size.

Between building our own products and bringing on teams from our acquisitions, Mailchimp has built a prolific engineering team. Like any startup, Mailchimp’s early days were scrappy—in the 2000s, the team was small but mighty. As the business grew, so did the needs of the team: in 2017, we went from a DevOps model to a full engineering model, and today, Mailchimp boasts over 400 engineers across multiple teams.

But with a nearly 20-year history, we’ve certainly seen some bumps in the road along the way. While the complexities of the business changed and the engineering team grew, we didn’t always turn to our own developers before making critical decisions about developer products. Just look at when we changed our Transactional pricing model in 2016, which forced Transactional users to sign up for Mailchimp to continue using the product. This switch didn’t make sense for a lot of Transactional users, and in the process, we eroded developer trust.

That erosion in trust taught us that we need to\_ *[*listen hard*](https://mailchimp.com/culture/how-failure-is-a-part-of-mailchimps-dna/)* \_to developers, and turn to our own team before making sweeping changes. That humbling lesson kicked off a series of changes that led us to where we are today—a brighter day for our developers. We’ve spent the last few years holistically researching the developer community, investing in resources like our pretty dope (if I do say so myself) developer-focused website, overhauled documentation, and [open source projects](https://mailchimp.com/mailchimp-acquires-reaction-commerce/).

But listening hard to developers also means listening hard to our own engineers. What kind of work excites them? What keeps them happy? How can we ensure we have a safe, inclusive work environment to attract and retain the most diverse and talented engineering team? While we’re working on some audacious plans [to solve for that internally](https://mailchimp.com/commitments/?ref=newsroom), I’ve been brought in, in part, to harness the incredible talent and energy on our engineering team and share some of their brilliance with the world.

### Supporting a growing developer community

With this group of outstanding engineers helping to pave the way, Mailchimp is investing in the tools, resources, and humans needed to support and grow a healthy developer community. I’m truly honored to be leading the charge as Mailchimp’s first developer relations person. Here’s how I approach DevRel, and how I’d like to translate that to Mailchimp’s developer programs and community.

First, I’d like to see us **elevating developer stories** a lot more than we do now. Our internal engineers solve complicated and interesting problems, day in and day out, and our partner developers are building cool shit. We want to shout this from the rooftops and start celebrating these wins!

We’re also hoping to share tactical advice, investing in more guides, tighter docs, relevant API and product updates, more sample apps, and more content types throughout /developer—can you say video? Can you say 3D hologram? OK, maybe no 3D holograms anytime soon, but we’re letting ourselves dream! The goal of this more tactical content will be to make building with and on the Mailchimp platform as\*\* frictionless \*\*as possible for developers.

And while we’ve always wanted our developer experience to be as smooth as possible, right now we’re especially invested in opening up more\*\* possibilities \*\*to build on top of Mailchimp. We have 11 million active customers, and try as we might, we can’t solve for every one of them. We want to make it as easy as possible to build the next great app that will empower our customers in ways we couldn’t have imagined. By working towards surfacing more easy-to-use, well-documented APIs, we want to help unlock developers’ creativity to build something great.

Another way to move toward a less friction-ful, more inspiring building experience is to start making room for **community and connection**. Historically, there haven’t been any dedicated spaces for developers working on and with Mailchimp to connect with each other: to share ideas and feedback, to help each other work through issues, to build networks, and hell, maybe even to make some friends? I want to change that. These efforts will begin in earnest in 2022, but know that I’m working on the inside to get y’all connected. And now that we’re looking forward to a post-Covid world…maybe we can even meet \_in person \_sometime! Freddie stickers are in your future, folks.

Ultimately, at Mailchimp, we want to empower our developers to empower the underdog, while ensuring we’re building a community that’s inclusive to any developer who’d like to join us in that mission. We want to hear from you. What’s missing from your experience building on Mailchimp? What’s stopping you from building the next great marketing app to support our millions of customers? I want to build these programs and this community around your needs, so get in touch anytime: [@SarahJaneMorris](https://twitter.com/sarahjanemorris), [devrel@mailchimp.com](mailto:devrel@mailchimp.com)

### Introducing “Mailchimp Engineering”

So here we are, on one of the first outward-facing steps towards helping developers empower the underdog: our brand-new blog, “Mailchimp Engineering.” We wanted to give our engineers the space to talk in public about those complicated and interesting questions they tackle daily: the unique types of problems they have to solve and the creative ways that they solve them.

On this blog, you can expect to read some powerful longform posts on software development—problems large and small, and how we approach them here at Mailchimp. We’ll be starting with tales from our own engineers, but eventually we’ll share experiences from the broader community as well, including external developers building apps on Mailchimp, or using Mailchimp Transactional or Open Commerce.

I’ll close with a teaser or two: while 2021 is all about foundations, we have some big things coming in 2022, things you’ll want to be ready for. So stay with us, tune into this blog, and watch as /developer evolves into a complete Mailchimp Developer destination. And in the meantime, keep your eyes peeled—we’re planning on posting something epic here monthly. We’ve been working with actual tech journalists and editors, and I’m so proud of the posts our engineers have in the queue. Can’t wait to go on this journey with all of you!