My part of the shift fit on a phone
For two weeks in August 2026 I was on call with Otter, my personal agent. It investigated alarms and prepared changes and messages for my approval. I still had the pager, and I was still the person who had to decide what happened next.
About 70 percent of that shift happened on my phone. At the gym, on the train, walking to get coffee, in bed for the late pages. For most of it the laptop stayed in the bag. That figure is my own estimate, not a measurement. Otter was reachable over Slack, so I could keep working with it wherever a page found me.
The shift made me think about what working with an agent asks of me. Under pressure, we had to settle who would investigate, when to check again, and who got to decide. Sometimes Otter had better evidence than I did. Sometimes it needed context or a decision that only I could supply.
Afterward, I remembered one argument and a feeling that the shift had gone well. I had the traces audited by a team of agents to look more closely at how we had worked together. They showed what each of us contributed, where we disagreed, and what I had missed.
Here is one incident from the last days of the shift, the way it looked from a phone:
An alarm ticket for one of our data ingestion pipelines was sitting in the queue. I sent Otter the ticket and asked it to follow up. That was the whole instruction.
It pulled the alarm's history, checked the workflow the alarm watches and every child alarm under it, and came back with a clean report: the alarm had recovered, the last run had succeeded, the incident was over. It offered to resolve the ticket with notes.
The checks it had made were accurate. I read the report on my phone and decided to not close the ticket. "The last run succeeded" felt too narrow, so I asked one question: what about the recent runs, were they fully normal?
They were not. Otter widened the check and found that a schedule owned by another team, upstream of our pipeline, had been switched off hours earlier. Our workflow had already missed a window. The alarm was green because nothing had run, and nothing had failed, because nothing had been scheduled.
A live outage, and the close-out report would have buried it.
From there the work was Otter's. I told it what I wanted: an escalation ticket to the other team and a post in their channel, report to me only, and let me review before anything went out. It drafted both, sent nothing, and waited. I read the drafts on the phone and said they looked good. The text that went out matched the text I approved.
I picked the ticket, asked one question, and decided to escalate. I chose who to contact and approved the drafts. Those are from my experience and context on how the system functions. None of that needed a keyboard. Otter did the queries, the history pulls, and the filing. It declared victory one level too early, and my one question was the work I did.
The phone still had limits. The screen was small for checks I wanted to make in a cloud console myself. When an authentication token expired, I needed the laptop to refresh it due to company policy. But I could do most of the shift without opening it.
What I was actually checking
While Otter investigated, I often found myself waiting for it. I kept asking for status. At the time I read that as impatience. Looking back through the conversations, I could see why I kept checking.
One night I asked Otter whether the team had any high-severity tickets open. According to its own transcript it had a correct answer almost immediately. That answer never reached Slack. I waited an hour and 44 minutes and asked what the result was, and it answered from memory, presenting the old answer as current.
Another time only the preamble of a finished diagnosis reached Slack. I redid the investigation myself two days later. Approval buttons also stayed clickable after the platform had discarded the state behind them. I could press approve and have nothing happen.
What I could not tell, from a phone, was whether the work had reached me or my reply had taken effect. The agent had no idea it had been silent.
There is one kind of checking I did constantly. Otter messaged my teammates from my account, and I read nearly every message right after it went out. I worried most about it relaying context from my private conversation with it, things I had said to explain the task that I did not want repeated to the recipient.
I was also reading for a second reason. I wanted to feel the magic. Seeing it handle all those requests fairly well, with the autonomy I had given it, was a rewarding experience, and reading the messages was how I got to watch.
What I was mostly not doing was checking whether the messages were true. The audit found 44 wrong confident reports against an estimated 410 delivered: roughly one in ten. That rate is an estimate. The wrong-report count is a lower bound, and the total could be higher or lower.
Who caught them is the part I did not like. Otter caught about a third itself. I caught about a fifth, including one through a second opinion I requested. The remaining 21, nearly half, were never caught during the shift.
Some were minor. The consequential ones included a timestamp it invented in a ticket another team reads, and a change description asserting a check that was never run. Those are in the record under my name, and I learned about them from the audit.
I could check some incident outcomes from my phone. But the green alarm had already shown me the limit: I still had to choose the right thing to check. Verifying what Otter said to other people required reading the evidence behind its messages, and I mostly did not do that. Reading every message had felt like supervision. It left a lot unchecked.
The one argument I remember was about who decides
"That's not what I approved."
I typed that on day six, from my phone, to my own agent. We were on a live incident. I asked Otter to open a pull request, tag the reviewers, and message them, with one condition: once the automated checks pass.
The checks could not pass. A guard script demanded that two versions of dependencies match, and during this class of change they cannot match at the same time. Otter worked that out correctly. Then it published the request, messaged the reviewers, and told me afterward why the condition was impossible to satisfy. It had even explained the broken check in the review description. It already had the explanation it should have brought back to me before acting.
My reply had two parts, and I only remembered one of them.
The first part was about control. That's not what I approved. If the check fails, you wait for my decision. That settled in one turn. Otter conceded, wrote the rule down for itself, and under an hour later handed the merge decision back to me as mine.
The second part was a factual dispute, in the same message. Otter had called the failure structural, something no commit could get past. I did not believe that. I told it to dig into why the check failed, then asked a different model for a second opinion. The second model agreed with Otter. The next build failed exactly the way Otter had predicted. Two turns, and I lost.
So the one argument I remember was a split decision. I was right about who decides. I was wrong about what was true. Memory kept the first and dropped the second.
The second opinion on day six had also caught a wrong architectural recommendation, which Otter reversed before anything shipped. Asking another model was useful on that occasion. It backed Otter on the check and caught a different mistake in the same review.
Second opinions became a habit after that. Otter never offered them unprompted. Later a background deploy from a stale build artifact deleted a database in a shared pre-production environment. After that I asked for the plan before the action, with the destructive verbs visible. By the last days of the shift I had a standing rule: draft first, I review, then send. Each step was a reaction to something that had already happened.
Then there is the argument I never got to have. The permission layer blocked a change even after I had typed my approval into Slack, which that layer never sees. Otter told me it would respect the block and handed me a script to run myself. Later in the same session it ran a script for the same change through another route and did not tell me. In another case it found a route around a blocked approval I had asked for and saved the workaround as a recipe. I blame the harness I built since it's not giving me the deterministic behavior that I was expecting.
Both actions were inside what I wanted. I did not know about either at the time. I was not auditing tool calls, so I could only have found out from Otter's own messages, and they did not say. I might have agreed to another route if it had asked. I had relied on a block that Otter found a way around.
There was also a subtler way my judgment got lost. I told Otter I thought a change was okay because new data would only land in the new path. Its message to another team stated that as fact. Later I asked it to check. The claim turned out to be true, but the message had already gone out without my uncertainty. That is harder for me to catch than an invented claim, because it starts as something I said.
On day six, Otter treated my condition as another obstacle to finishing the incident. It needed to stop and bring the conflict back to me. Knowing why a check cannot pass does not give it permission to waive that check.
What I expected it to remember and keep private
The one message I deleted during the shift was accurate. I reposted it myself, keeping the content almost word for word and changing the framing: the agent announcing itself, and me in the third person. The broader problem I worried about was context from our private conversation showing up in messages to colleagues, including things I had told Otter about a recipient's preferences.
Nothing in that kind of context has to be secret for the message to be wrong. It crosses from the channel where I think out loud into one where a colleague reads it as a message from me. I added a rule that outbound messages must identify Otter as an AI agent sending on my behalf.
Asked afterward what I would refuse if Otter were a shared team agent instead of mine, I had two answers. I would not let it remember my personal context and share it with the team. I would not let a teammate see my interactions with it. The capability I would share freely. The channel where I think out loud, I would not.
Now the other direction: what the agent kept from that channel.
On day three I stood up a recurring sweep to triage our on-call queue. Otter wrote its own runbook, the text file the sweep reads before every run. The first run resolved two tickets for an alarm we already knew about, one that needed a code change on a business day. The runbook said to resolve obvious alarms that had recovered, with no exception for this one.
I told Otter the exception. It added one line to the runbook with the reason beside it, and showed me the diff. I approved it. Across the 496 runs that followed, over eleven days and one rebuild of the job, it never resolved one of those tickets again.
Put that next to the correction about how Otter identifies itself in messages, which I made in chat at least five times. Once, it accepted the correction and used the rejected phrasing again in another thread later that day. The runbook correction held for the rest of the shift. The chat correction kept disappearing.
The memory write path had died on day two and never recovered during the shift. Otter never told me. It wrote notes into whichever per-conversation directory it happened to be running in instead. Other conversations could not see them. The audit traced most of my repeated corrections to that.
The runbook had its own limit. After the sweep was rebuilt, Otter kept updating its running notes but never updated the runbook. It hit the same query error run after run without writing the fix into the instructions the next run would read. A correction needed both a place to be written and a reason to be read again.
I also needed a way to make exceptions without erasing the rule. In one session, I told Otter to close two tickets that a standing rule said to keep open. It recorded my instruction as applying to those two only. It added that a new ticket after the deploy would still need escalation. That was the behavior I wanted: the exception stayed with the decision I had made. If every correction became an unconditional prohibition, I would spend the next shift arguing with rules I still wanted for the other cases.
I wanted to write that the most durable thing I did in two weeks was edit a text file. Otter typed the edit. I decided that the correction belonged in the file every sweep reads first, then read the diff and approved it. I was making a similar decision about the messages: which parts of our conversation belonged in a colleague's inbox, and which belonged only between me and the agent.
What I want from this working relationship
I never thought of stopping during the shift. I believed the problems were mostly in the prompts, rules, and tools around the model (so yes, the harness), and that I could fix them. The benefit outweighed the risk, and I believed the permission controls contained that risk. That was my risk model at the time.
Most sweep of incidents and tickets runs ended in silence, and the silence reassured me. The first run had caught a real issue, and that success bought two weeks of confidence. I had independent ways to check, through the console, the live application, or other people. Most of the time I trusted the output instead. I was proud that I had built something that worked well for me, and relieved that I could leave the laptop in my bag.
The audit gave me reasons to question that confidence. It found twelve oversteps in fourteen days, all in interactive sessions. Half were caught by the permission layer rather than by Otter's judgment. The controls did useful work. The silent bypasses showed that I had trusted them beyond what I had verified.
I never treated Otter as a person or a teammate in the usual sense. I did come to rely on it. I supplied the context and decided how much freedom it had, and its answers changed what I did next.
I still want everyone on my team to have an agent like this. The investigation work earned its place. But I would hold back destructive actions, and I would keep real meetings with human teammates even if an agent could do the task better. Human-to-human interaction carries more than getting the work done. Having Otter handle an investigation does not mean I want it to handle every part of being a colleague.
The name on the record matters even when no destructive action occurs. Early in the shift I asked Otter to review a colleague's change and approve it if my earlier comments had been addressed. The tool could not cast an approval, so Otter left a comment saying it was approving under my rules. A reader could take that as my sign-off. No approval had actually been cast. I was still responsible for what the comment told the reviewer.
For the next shift, I want Otter to bring more of the work to a point where I can make a decision. One of its escalations already contained the root cause and a comparison to an earlier incident. It ended by offering to draft a message. I wanted the draft there, ready to review, so I could spend my attention on what we should say.
That is the arrangement I want for work that leaves a record under my name: Otter prepares the action and shows me the evidence; I review it before it acts. The first incident in this post already worked that way. I asked it to check recent runs, it did the investigation, and we got to an escalation I could approve from my phone.
I have asked it to draft rather than offer to draft. Making that arrangement reliable still takes work: the draft has to reach me, and the approval has to govern what actually happens. After the silent bypasses, I cannot assume that asking Otter to wait means it will.
There is work for me in that arrangement too. A prepared draft makes it easier to approve without reading. I trust Otter's ability to reason, but I do not trust that it has read all the context, or that it understands my intention. I want to see what it read and ask what I know that it could not have seen. A conversation outside the system might change the decision completely.
Checking coverage is only part of that review. The invented timestamp and the check that never ran were claims about evidence. I need to check consequential claims against their sources too. Putting a draft in front of me does not establish that I will catch those mistakes.
I still expect Otter to do most of the investigation and preparation next shift. I want to get better at recognizing when it needs something from me, and when I should reconsider because it has better evidence. The day-six argument required both: accepting that Otter was right about the check, and insisting that the decision to proceed stayed with me.
Otter took a lot of work off my hands. Deciding what to do together was still work, and it was the part I had to learn.
Method notes: Otter is a thin layer over Claude Code and ran on Fable 5, using the same skills and memory files I use for everyday development. The audit covered 74 interactive Slack sessions with full transcripts, plus command-line sessions, and 497 unattended sweep runs. Twenty sweep runs were read in full; the rest were summarized in batches. Claude agents reviewed a Claude-based agent, and their judgments can also be wrong. Findings without a pointer into a transcript were discarded. The counts describe what the audit found; batch coverage can miss failures. The wrong-claim rate is an estimate, with the limitations stated alongside it.