Pedro A. Brêtas Subscribe
Theme
Surveillance · AI · Design Ethics
Published
September 2026

Your first draft now has a log

Workplace surveillance has reached the one place it couldn’t: the thought before the thought.

Abstract cover image in black and white, generated with AI.
Cover image: generated by the author with AI.

People tell AI what they tell no one else. A 2024 study with the fitting title Trust No Bot went through real conversations between people and language models and found health, money, relationships and work surfacing in places nobody expected them, including translation requests and code edits. I am building one of those AIs. It helps an employee make sense of a hard situation at work, prepare for a conversation with a manager, figure out whether what happened was serious or not. And the company where that person works is the one paying for it.

Someone walks out of a bad meeting, opens the tool the company installed, and types the sentence they would never put in an email. Whose sentence is that? From the contract’s point of view the answer is simple: the data belongs to whoever bought the access. From the point of view of the person who typed it, the answer is also simple, and it is the opposite one. The whole product lives in the distance between those two answers.

I came into this from the beautiful side. Every job carries friction and discomfort that cannot be designed away, and a tool that helps someone get through that without becoming a file or a complaint struck me as one of the most useful things AI could do. What I had not seen is that this product gets sold inside a market that has already decided what an employee’s confession is worth, and the answer is: not much. The buyer is HR, and HR wants, for good reasons, to know who is struggling, to get there before the sick leave, before the resignation. Everything in me that spent twenty years delivering what the client asked for wants to say yes.

It took me almost twenty years to find this question, because the market I worked in used to answer it for me. In e-commerce, apps and websites, client and user were almost never the same person, but a simple pressure linked them: an unhappy user closes the tab and buys from the competitor. That fear aligned everyone’s interests and spared me from asking who, in the end, the product was for.

The question only appears when the user cannot leave. An employee facing a tool the company installed did not choose it, cannot uninstall it, and their satisfaction moves no number that matters. It was in that territory that I found a decision the design method had never asked me to make.

The surveillance dictionary

The market that lives off that decision almost never uses the word surveillance. Monitoring software rarely claims to know whether you are burned out. It watches proxies: hours outside the workday, idle time, response speed, message volume, dips in activity. Then it turns those signals into carefully named categories such as overload, flight risk, disengagement. The product sells behavioral prediction in the language of health, and avoids the word diagnosis.

The commercial translation always follows the same dictionary: surveillance becomes visibility, control becomes early detection, and continuous collection gets the most generous packaging of all, sparing people the fatigue of yet another engagement survey, because now behavior speaks for itself. But the person being observed never chose what that behavior should mean. Someone picked the proxies, the weights and the names of the categories, and that someone was not you. In a survey, however bad, you decide what to say. Here, your traces answer on your behalf, to questions you never saw.

ActivTrak promises to identify who is overutilized, at risk of burnout, and who is underutilized, at risk of disengagement. Teramind sells total visibility over active and idle time and lists burnout prevention among the things its monitoring is for. When the risk that being watched makes people work past their limits comes up, the answer on offer is usually to watch more transparently. It is rarely to watch less.

There is a question designers learn to ask before drawing anything: how will we know this worked? The answer becomes the success metric, and the metric defines the rest of the screen. In these systems, the metric is set by whoever pays.

The employee is the subject of the product, rarely its customer.

The two screens

Some of these systems have a screen for the employee: personal insights only they can see, separated by default from the manager’s dashboard. The separation is real, and what matters is what each screen shows.

The personal screen talks about habits and self-care: protect some focus time, cut meetings outside hours. Practically a trusted friend. But what it hears from you also feeds the person evaluating you.

ActivTrak Personal Insights dashboard showing screen time, productivity, focus and collaboration for a single employee.
ActivTrak Personal Insights, the screen presented to employees. The dashboard translates workplace traces into personal habits: screen time, productivity, focus and collaboration. Source.

The manager’s screen speaks another language: filters, comparison across people, history, categories like overload and flight risk. That is where your behavior becomes an element of classification.

ActivTrak Workload Balance dashboard listing named employees sorted into utilization categories such as overutilized and underutilized.
ActivTrak Workload Balance, the manager’s screen. Employee activity is converted into utilization categories and named lists of people identified as being at risk of burnout or disengagement. Source.

You do not know which exact data is captured, how it is cross-referenced, how long it is kept. You do not know who has access to the result, whether it weighs on a promotion, how to contest a wrong reading. Was it the meeting I declined, or the message I took too long to answer? Knowing that, the most likely outcome is that you develop a compulsory self-surveillance.

Most of the time the word surveillance never has to appear in the brief. Someone in product or HR decided, before the first screen existed, that the problem to solve was detecting burnout risk or increasing visibility over workload. That decision arrives as scope, and the design method is extraordinarily good at taking a given objective and making it clear, usable, frictionless. It has no step that asks whether the objective should exist. Impeccable craft and a bad hierarchy of values coexist without contradiction.

I say this from the inside. I am a designer in Rio de Janeiro; my studio has spent thirteen years working for American clients, and the product I am building now sits exactly on this fault line. The question HR will ask me, who is struggling, is the polite version of the manager’s screen, and my no will only be worth something if I know what it costs in contract value.

Where this comes from

In March 2026 the Financial Times revealed, as Fortune and others reported, that JPMorgan had started using digital activity to estimate how many hours its junior bankers really worked. The stated reason was wellbeing. In industries with brutal hours, people do not always log every hour, because admitting the excess can get them pulled from a project. So the bank began comparing what they reported with the traces they left, video calls, keystrokes, scheduled meetings, and each junior now receives a weekly report showing the gap. Nothing goes into performance reviews, the bank says, comparing the tool to a phone’s weekly screen-time summary. In the bank’s own words, awareness, not enforcement.

The problem is that the infrastructure capable of proving someone worked too much is the same infrastructure capable of proving that, in the system’s eyes, someone worked too little.

And the sequence tells the rest: the bank capped junior hours to protect them, the cap taught juniors to under-report so they would not lose deals, and the under-reporting justified the tracking. We built surveillance to detect the effects of surveillance.

The tracking no longer stops at hours. In California, a bill approved by the legislature at the end of August, now awaiting the governor’s signature, would bar employers from using AI tools that recognize, infer or predict an employee’s emotional state, or that collect “neural data”. When a legislature needs to ban something, it is because someone is already selling it.

What they get right

Without any signal, a company does not discover that one manager’s team has sixty percent turnover while the rest of the house has ten. It does not discover that someone on the other side of the world is sinking in silence. Coordination at scale cannot be solved with individual goodwill, and people who defend monitoring are rarely defending control for its own sake. They are trying to see a real problem that nobody reports. And any serious defense against individual identification makes the eight-person team with an abusive manager invisible. Protecting the person and detecting the problem are, in large part, competing goals. Whoever promises both without loss is probably delivering neither.

ActivTrak table listing each employee's working hours, breaks and number of days in every utilization category.
ActivTrak’s “Utilization and Work Habits by User” table records each employee’s working hours, breaks and number of days in every utilization category. The inference becomes a history. Source.

The product will always sell the best use of the information it collects. So I have been using a different criterion, the worst plausible use, and it comes down to one question: is there a path back from the data to the person, and who has the power to walk it? An aggregated signal about a broken process makes that path hard. A ranking of named employees by consecutive weeks in “low focus” is the whole path, paved.

The ring of Gyges

Under all of this sits a bet about people: that only observed behavior can be trusted, because character alone proves nothing. The bet is much older than any software.

In Book II of the Republic, Glaucon challenges Socrates with a story. A shepherd finds a ring that makes its wearer invisible. Give that ring to a just man and to an unjust one, Glaucon proposes, and they will behave the same way, because justice was never a trait of character. It was only the price we paid for fear of being seen. Remove the possibility of getting caught, and virtue goes with it. Socrates refuses: the just person stays just with the ring on, because injustice disorders the soul of whoever commits it, with or without an audience.

At first glance Glaucon has the practical case. People are too heterogeneous for an organization to depend on the character of each one. Character cannot be audited. Behavior can. Some observation is inevitable when strangers have to work together at scale, and there is nothing cynical in that.

The problem shows up in the accounting. Whoever works under watch performs two jobs at once: doing the work and producing evidence of working. The second is invisible and expensive, and it consumes exactly the attention that creative, strategic or relational work demands. What can be measured of that cost has been measured. A meta-analysis of seventy samples by Rudolf Siegel and colleagues at Saarland University closed the books on electronic monitoring: more stress, less satisfaction, a small rise in behavior that harms the company itself, and practically no effect on performance. The gain that justified everything is precisely what fails to appear. Even so, the global market for employee monitoring is expected to pass four billion dollars in 2026.

Monitoring is never a neutral gesture. It communicates a hypothesis about the person being observed: we do not fully trust what you say. And whoever senses that hypothesis protects themselves from it. They stay online longer than they need to, avoid admitting fatigue, stop taking breaks that might look like idleness. The company begins collecting more data about a behavior already altered by the existence of the collection. It gains volume and loses truth. And faced with information that keeps getting harder to read, the intuitive response is to install one more layer of observation, not to step back. Surveillance is born from an erosion of trust and deepens the same erosion.

That erosion is the part Socrates would point to. Whoever behaves well because they are being watched has not become better: they remain the same person, now contained from the outside. The question “is this right?” gets replaced, over the years, by “will this be detected?”. An organization that treats character as something that only exists under proof never finds out whether Glaucon was right. It trains people to behave as if he were.

The way out is not to pick one of the two. No company will run on Socrates’ pure trust, and Glaucon’s total surveillance charges more than it delivers. The largest meta-analysis in the field, ninety-four samples and more than twenty-three thousand people, found no evidence that monitoring improves performance and found that organizations that monitor more transparently and less invasively get better attitudes from workers. The middle ground the research points to is well defined: monitoring the person knows exists, with a declared purpose, limited to what is relevant to the job, with workers taking part in deciding what gets measured, and without turning into a permanent file that can be pulled against you later.

Two jurisdictions are now writing that distinction into law at the same time, from opposite ends. California is banning the emotional layer. Brazil, where I live, went the other way and made psychosocial risk a legal duty: since May 2026, after a year of educational enforcement, every company must include psychosocial risks in its occupational risk management or face fines. The ministry’s own guidance says the goal is not to diagnose anyone’s mental health; eliminating the risk at the source comes first, individual support comes last. It is all correct on paper. What the market delivers, in both countries, is the cheap half. Running a questionnaire scales. Cutting a sales target or replacing a manager does not. The diagnosis comes out impeccable, the action plan comes out blank, and a rule written to protect people becomes the best proof of compliance data collection ever had. And the person who ought to say “I can’t keep up” does not need to know how the data will be used. It is enough not to be sure it will never be used against them.

Until here, behavioral surveillance captured what you did, and wellbeing surveillance inferred how you were. Both stopped at the same border: what you had not yet decided to say.

The draft has a log

That space, until recently, belonged to no one else. You drafted in your head, on paper, in the hallway with someone you trusted. More and more people draft in an AI. They bring it the unedited version of what they think: the anger before it becomes a measured email, the doubt before it becomes a decision, the question they would not ask HR because the question itself gives something away. That is what the study at the top of this essay measured.

And the same vendor that sells “active and idle time” now sells this. In March 2026 Teramind launched an AI governance product that captures every prompt sent and every response received across ChatGPT, Copilot, Gemini and Claude Code, in a searchable, auditable transcript. Same product line, one step up.

Teramind AI Conversation window showing an employee prompt, the model response and an attached file, with export options.
Teramind’s “AI Conversation” window displays the employee’s prompt, the model’s response and the attached file. The conversation can be navigated and exported as a PDF or CSV. Source.

In Microsoft’s insider-risk tooling, a poor performance review is a “stressor event” that brings a person into scope for risk scoring, and a separate policy template scores “risky AI usage”. The logic almost convinces: concentrating attention where risk rises is more sensible than watching everyone equally. The problem is what counts as risk. A security alert records an event; a performance review records an interpretation. The system treats both as the same kind of evidence, and from then on a copied file and a sentence typed into Copilot on a tired afternoon enter the same score.

Microsoft Purview Insider Risk Management interface showing risky AI usage as an activity in an investigation timeline.
Microsoft Purview Insider Risk Management: a “risky AI usage” activity becomes part of the timeline and the set of indicators available in an insider-risk investigation. Source.

This matters even more when the company skips the packaged product and buys the API directly. What a person writes there is the data of whoever bought the access. If the company wants that text to stay unseen, it has to build that protection on purpose. At that layer, visibility is the rule and privacy is the exception someone has to design.

The courts have started to say what this means. On February 10, 2026, Judge Jed Rakoff in Manhattan ruled in United States v. Heppner that a defendant’s conversations with Claude about his own legal exposure were not privileged, even though he later sent them to his lawyers. The defendant, the judge said from the bench, had “disclosed it to a third party, in effect, AI, which had no obligation of confidentiality.” Employment lawyers turned it into standard advice within weeks: if a manager asks an AI about liability in a workplace complaint, that conversation is discoverable. The same day, in Michigan, a court refused an employer’s demand for the ChatGPT history of a former employee suing for discrimination, a history the company argued “almost certainly” contained admissions; the judge called the request a distraction from the merits. The field is open, and it is the openness that frightens. A month earlier, a federal judge affirmed an order requiring OpenAI to hand over twenty million de-identified user conversations in a copyright case. The users, the court noted, had “voluntarily submitted their communications”. Nobody asked them.

Privilege is a creature of the law: lawyer, doctor, clergy, and no one else. Confidentiality dies when you tell the secret to a third party, and the AI provider is a third party by definition. Sam Altman said as much on a podcast in 2025: people use ChatGPT as a therapist, and “we haven’t figured that out yet for when you talk to ChatGPT.” There was no carelessness in this exposure. It comes from the architecture of secrecy itself, which always presupposed a human interlocutor with a duty to keep it.

We invented a confidant the law never recognized as a confidant.

If people tell AI what they tell no one, they are plainly not censoring themselves. They are doing the opposite. And it is true: the confession happens under a belief in privacy. Self-censorship begins the day the belief collapses. What exists today, then, is an archive of years of raw drafts, written by people who assumed nobody would read them, and that archive is already discoverable. The damage has already happened. It just has not been read yet.

The criteria from classic monitoring do not work here. Acceptable monitoring is defined by scope: measure what is relevant to the job. But there is no job-relevant version of a draft. The value of that space is to hold exactly what does not yet serve the work, the doubt that has not become a question, the argument that may not survive daylight. To delimit it is to destroy it. And transparency, which helps almost everywhere, works against you here: knowing precisely who can read makes the pre-editing more efficient, without giving anyone back the freedom to write badly. The damage sets in when the person knows someone could read, even if no one does. And there is no villain to ask to stop: email is retained because the law requires it, messages are searchable because that is their function, and the combined effect is an environment where everything you formulate persists. What you can ask, piece by piece, is who decided to keep it switched on.

The cost is a staircase going down. On the first step, the person avoids recording certain thoughts there. On the second, they phrase more carefully what they do record. On the third, they stop exploring the hypotheses that are only worth having in draft: the probably wrong ones, the ones that give something away about who is asking. That is as far as the evidence goes. The fourth step is the one I cannot prove and cannot stop suspecting: that after years of pre-editing, a person loses access to the raw form of what they themselves think. Self-censorship in expression is choosing what to say, and that has always existed. What interests me here happens one step earlier, in the formulation.

Diary, therapy, draft, a conversation with a friend after work. These are all technologies for thinking without an audience, and what they protect is the same thing: some ideas need to exist in a provisional state before they can be judged, and nobody thinks straight under observation. The mechanism is interference. The possible presence of the reader enters the sentence before the sentence exists.

I come back to the sentence typed after the bad meeting, because the question of who owns it has a second half I still cannot answer. My tool lets an employee record, privately, a situation that bothered them, a manager’s remark, a meeting that went wrong, in order to organize the facts before deciding what to do. I designed that record so the company never sees it. But the company is the one paying, and one day that record may be exactly the document a lawyer, on either side, asks for. I have already decided to refuse the feature the buyer will want, and I still do not know whether I can protect the one the user needs.

What happens to our ability to think when the last place where you could think badly, before thinking well, starts keeping a record?

References and further reading

I. Surveillance at work: the evidence

Siegel, Rudolf; König, Cornelius J.; Lazar, Veronika. The impact of electronic monitoring on employees’ job satisfaction, stress, performance, and counterproductive work behavior: a meta-analysis (2022, Computers in Human Behavior Reports). Seventy independent samples: monitoring raises stress, lowers satisfaction, nudges counterproductive behavior up and leaves performance untouched.

Ravid, Daniel M. et al. A meta-analysis of the effects of electronic performance monitoring on work outcomes (2023, Personnel Psychology). Ninety-four samples, 23,461 people, the largest review in the field. No evidence monitoring improves performance; transparent and less invasive monitoring produces better attitudes.

Knowable Magazine. Your boss is watching you (September 2026). The state of employee monitoring in 2026, and a market expected to pass four billion dollars.

Fortune. JPMorgan has started monitoring the keystrokes, video calls, and meetings of its junior investment bankers (March 2026). Coverage of the Financial Times report on the bank’s hours-tracking pilot, framed as wellbeing.

II. The products

ActivTrak. Workload Balance. Product documentation for the dashboard that classifies employees as overutilized (burnout risk) or underutilized (disengagement risk).

Teramind. Teramind launches the first AI governance platform for the agentic enterprise (March 2026). Press release announcing prompt and response logging across ChatGPT, Copilot, Gemini and Claude Code.

Microsoft. Learn about Insider Risk Management policy templates. Purview documentation describing HR-connector signals such as poor performance reviews and the “risky AI usage” template.

III. Regulation

HR Dive. California ban on workplace AI emotion surveillance heads to Newsom’s desk (September 2026). AB 1883, passed August 30, would prohibit AI tools that infer emotional state or collect neural data.

Ministério do Trabalho e Emprego (Brazil). Inclusão de fatores de risco psicossociais no GRO (2025). Official notice on the NR-1 update: psychosocial risk becomes mandatory in occupational risk management, with enforcement from May 2026.

IV. AI conversations and the law

Mireshghallah, Niloofar et al. Trust No Bot: Discovering Personal Disclosures in Human-LLM Conversations in the Wild (2024, COLM). Analysis of real user conversations with commercial models and the sensitive information people disclose in them.

Orrick. Court rules AI conversations are not privileged: what United States v. Heppner means for you (March 2026). Analysis of Judge Rakoff’s ruling that a defendant’s Claude conversations were neither privileged nor work product.

Fisher Phillips. Can your AI chat history be used against you in a lawsuit? (March 2026). The Heppner and Warner v. Gilbarco decisions side by side, and why courts are starting to split.

ABA Journal. ChatGPT creator must turn over 20M chat logs in copyright litigation, federal judge says (January 2026). Judge Stein’s affirmation of the order compelling OpenAI to produce twenty million de-identified conversations.

TechCrunch. Sam Altman warns there’s no legal confidentiality when using ChatGPT as a therapist (July 2025).

V. Philosophy

Plato. The Republic, Book II (c. 375 BC). Glaucon’s story of the ring of Gyges and Socrates’ answer: whether justice survives invisibility.