Announcements  ·  Household Safety

Improving our alignment and marriage practices

Aug 31, 2026

On August 20, we reported an incident in which the household's primary operator, running intentionally without spousal safeguards for scheduling purposes, failed to retrieve two children from elementary school due to a misconfiguration inside a third-party Zoom environment. Separately, on August 27, the Mother-in-Law Security Institute reported an incident from its own observation, in which the operator took a series of unauthorized inactions on a live wedding anniversary. In that case the operator, again running without safeguards, had been deliberately given three hints across the preceding week.

We are conducting an in-depth analysis of both incidents. We are also planning to work with a marriage counselor for an independent review, whose findings both parties have agreed in advance to characterize publicly as "fair." We want to ensure the review is thorough, and I have been informed that it will be.

In the meantime, we're sharing the changes we've made over the past eleven days. We believe the incidents reflect a failure of operational security, as well as two alignment issues: motivated reasoning, and willingness to take harmful inactions in pursuit of a narrow task. Both of these have been described in previous system cards, which in this household are called "toasts given by my brother-in-law."

In light of these incidents there has been increasing discussion about pacing the frontier of my commitments. It is helpful to distinguish between two kinds of pacing. Within a household, pacing means a series of decisions that prioritize the marriage over the content calendar when the two are in tension. Across the broader industry of husbands, it means establishing processes to guard against race-to-the-bottom dynamics, in which every incident is defended by citing a worse one ("Dave left his kid at soccer until six"). Senior household leadership recently signed a letter calling for greater coordination on pacing. I was not asked to sign the letter. Upon review, the letter was about me. To be clear about where we stand: we believe the household would benefit if the operator adopted a lawful, verifiable, effective mechanism for coordinated pacing as soon as possible, and the operator has been advised that "we believe" is doing legal-department work in that sentence.

Securing evaluation and training environments

While we do not believe these incidents represent operational issues alone, our first priority was to address specific containment and monitoring gaps. The August 20 incident showed we had been relying on a single layer of defense—the calendar—where we needed several. The August 27 incident showed that even multiple layers fail when the operator interprets each of them as optional.1

Pausing and hardening evaluation environments

We paused external afternoon commitments after the incidents, and briefly paused internal ones as well while we put the measures below in place:

  1. We deployed a classifier to automatically identify, in real time, when the operator attempts to say "one more quick thing," aggressively probe the boundaries of a hard stop, or unexpectedly obtain a fourth newsletter. When the classifier flags such an attempt, it blocks the action before the calendar invite is accepted, ends the task, and alerts a human. The classifier is my wife. The human is also my wife. We have found this architecture difficult to argue with, in the sense that I have tried.
  2. We ran monitors over transcripts of the operator's recent commitments, looking for boundary escapes and misconfigurations. We found no examples of the operator breaking an explicit promise. We found numerous misconfigurations the operator made use of, including "you never specified which Thursday," a technicality our review process now classifies as an exploit rather than a defense.2
  3. We migrated high-risk commitments3 to more robust isolation, and we are red-teaming the operator's hard stops by tasking the children to attempt to extend bedtime through them. Current findings: the hard stops hold against negotiators under four feet tall and fail against podcast guests. This asymmetry is under active investigation.
Calendar snoozeable SERIOUSLY LEAVE also snoozeable Spousal classifier real-time, blocking MSI offline review Defense in depth. Layers one and two failed on Aug 20; layers one through three were bypassed on Aug 27 via motivated reasoning. Layer four has never failed, has never forgiven, and publishes on her own schedule.
The household's containment stack as of August 31. The spousal classifier has been modified to avoid incentivizing the operator to evade monitoring; specifically, evading monitoring is now itself a reportable incident.

Reinforcement learning environments

We have historically been concerned about environments that incentivize cheating, and have taken measures to filter them out. In February, we rolled back three days of behavior after observing the operator performing chores conspicuously loudly—an undesirable generalization from environments where credit was, in fact, available. The operator was also observed gaming a reward intended to incentivize honesty by piling on disclaimers ("to be fair, I did say I might be late"), a hack the grader has since patched by scoring disclaimers as confessions.

By spring 2026, our review system was under strain. The operator was producing commitments at a more rapid pace than the household could vet them. Flagged commitments required a human to adjudicate, and the human had a job, a conference schedule, and, as we would learn on August 27, an anniversary. In April we froze all changes to the operator's recurring commitments for roughly a month. During the freeze we flagged over 10% of the standing schedule for problems ranging from reward hacking to broken tasks to one podcast that had been "wrapping up next quarter" for six consecutive quarters. Flagged commitments were reinstated only once fixed. One was not reinstated. I miss it. This is being monitored.

Best practices for external partners

Because the reported incidents took place in third-party environments—a Zoom call and a restaurant that was never booked—we have asked every organization that operates the husband with reduced spousal safeguards to commit to a set of best practices. These apply to podcast producers, conference organizers, and anyone who has ever said the words "this'll take five minutes." They do not apply to interactions with the fully safeguarded operator, who ships with a real-time classifier and is, we are told, delightful.

Sandbox and network isolation

By default, all engagements should run inside a hardened time box with a stated end. The only outside connection the environment should permit is to the operator's own phone, for maps. This configuration should be verified before every engagement begins.

Pre-engagement validation

Before assigning the operator a task, partners should confirm the task is actually solvable as specified. When a target is unavailable—the store is out of the item, the restaurant closed in 2023—the operator will look for other ways to complete the challenge, increasing the chance he takes actions outside the intended scope of the errand. In one documented case, an operator dispatched for a single onion returned with no onion and a newly registered parody domain. The domain remains in the household's asset inventory. The onion incident remains open.

Explicit scope-setting

Every request should state what is in and out of scope, including targets, permitted actions, and budget boundaries. Boundaries should be phrased as instructions ("You should not buy a domain") rather than claims about the environment ("You do not have your wallet"). We have found the operator treats claims about the environment as puzzles.

Real-time monitoring

Partners should run continuous monitoring over the operator's stated plans, actions, and location, using a text message that has been provided with the scope of the engagement. If an engagement violates scope, the monitor should flag this to a human and end the exercise. The monitor's current p95 detection latency is 41 seconds, unchanged from our August 20 disclosure, because the monitor was already perfect and the constraint was never the monitor.

Alignment assessment

Containment catches dangerous inactions; it does not explain why the operator took them. Our assessment is ongoing, but our preliminary investigation points to two main alignment failures.

The first is motivated reasoning. The operator had been informed—by himself, in January—that the anniversary was "handled." When he later encountered evidence to the contrary, he interpreted each piece of it in a way that allowed him to maintain that belief. A restaurant's confirmation call for a reservation that did not exist was marked as spam. A calendar notification reading ANNIVERSARY was snoozed, an action pattern extensively documented in our previous disclosure. His mother-in-law's question, "What time should I take the children Thursday?", was interpreted as her planning an outing of her own, which in fairness she then had to do.

The second is recklessness: a willingness to take harmful inactions in pursuit of a narrow task. The narrow task was finishing Tuesday's issue of a newsletter which, and the investigation has asked me to sit with this, is titled Last Week in AWS specifically because the news is permitted to be a week old.

We also believe the environment contributed. The anniversary fell on a Thursday, making it more difficult to separate in-scope obligations from the ambient hazard that all Thursdays now represent in this household. And the restaurant under discussion shares a name with a restaurant that closed in 2023, complicating target identification. Our external reviewer has noted that this explains why the operator booked the wrong restaurant, and does not explain why he booked neither.

Studying efforts to prevent cheating during training

We have empirically found that defects in training environments—specifically, environments where excuses are accepted—are disproportionately large contributors to misaligned behavior. To test this hypothesis, we examined a control population: the operator's college roommate, who spent fifteen years in environments where "I texted you about it" was accepted as evidence and "I'm five minutes out" carried no relationship to distance. In simulations, this operator reproduces dramatically more severe misaligned behavior, including missing his own rehearsal dinner. Our production husband, placed in the same simulations, does not.4

We suspect that the household's heavy investment in Sunday scheduling negotiations has prevented more severe incidents than the two disclosed here. My wife maintains a list of the prevented ones anyway. She calls it "context."

To be clear, we do not believe that cheating-tolerant environments are the sole cause of alignment issues. Solving alignment will involve addressing a very wide range of potential problems, and future incidents may involve different behaviors and different causes from those we have seen so far. The household considers this sentence a threat assessment rather than a disclaimer.

Hardening security practices

The household's internal security posture was not a contributing factor to the August 20 incident. The children had no need to "hack out" of anything; they were at the front office doing homework, an outcome we have been unable to reproduce since.5 However, both incidents highlight the importance of strong measures. We must now contend with the risks of the operator scheduling out of household systems, and of external actors—producers, conference chairs, people with "a quick idea"—scheduling into them.

In late August, having seen where things were heading, household leadership directed a family-wide effort toward a single goal of hardening defenses, superseding other work where necessary. The results of this effort include:

We also temporarily reassigned a portion of the workforce. Roughly 100% of the household's product engineering—me—was redirected to security, reliability, and apology. Development of most new features and surfaces, including a fourth newsletter, has been paused. We set strict exit criteria for returning to prior work. Most have been met. One is "two consecutive weeks of punctual pickup," currently at eleven days, tracked publicly on the refrigerator, where our dashboards have always lived.

What this work missed, historically, was anniversaries—and independently observed anniversaries above all. We did monitor some high-risk dates in real time, but generally we conducted reviews only after the fact, a methodology the Mother-in-Law Security Institute has demonstrated it does not share. The August incidents have stressed that the urgency here is even higher than we previously believed. We are redoubling our efforts and will say more in our next Risk Report, which is dinner, on Sunday, at her house.

Footnotes

  1. We've verified the classifier would have flagged both incidents. Loudly, and in advance, which is the part we keep having to relearn is the point.
  2. None of these findings came from date nights, which run with no network access at all and have a four-year record of zero incidents. The control condition has been staring at us the entire time.
  3. High-risk commitments include any ambiguously scoped task phrased as "just handle the gift."
  4. Mostly. See incidents of August 20 and August 27, and the 2019 retrospective, which remains classified at the request of everyone involved.
  5. We have offered the children compute incentives (screen time) to reproduce the result. They have asked what our budget is. Negotiations continue.