On September 2, 2026, The Wall Street Journal reported that being judged a good worker by the AI productivity-monitoring systems companies are rolling out has become a real factor in how employees are evaluated. When AI scores how people work, is it measuring the quality of the work, or only the appearance of working? Once activity becomes the target of the score, people shape how they look to the system, and the metric stops reflecting the quality of the work. Monitoring helps only when employers decide, and disclose, what they measure and why.

01AI now scores the working day, and employees are adjusting to the score

The Journal piece by Callum Borchers starts from a plain observation about today's workplace: it now matters whether the AI system your employer installed classifies you as a good employee. From there it treats a measure of gamesmanship, the effort to look strong in the system's eyes, as a realistic response on the employee's side.

The spread of monitoring shows up in survey data as well. In the American Psychological Association's 2023 Work in America survey of 2,515 employed adults, 51% said their employer uses technology to monitor them while they work. For half of the workforce, being monitored is no longer an exception. It is a condition of the job.

The issue worth examining is not who plays the game well. It is what the AI is actually turning into a number: the quality of the work, or the look of being busy. If the score tracked quality, there would be little room to raise it through gamesmanship. The fact that gamesmanship is being discussed at all is evidence that the score picks up something other than quality.

02Once activity becomes the target, the metric stops reflecting the work

What happens when a measurable quantity becomes the target of evaluation has been described for half a century. In the 1970s the British economist Charles Goodhart observed that any statistical regularity tends to collapse once pressure is placed on it for control purposes. The anthropologist Marilyn Strathern generalized the point: when a measure becomes a target, it ceases to be a good measure. The social scientist Donald Campbell wrote that the more a quantitative social indicator is used for social decision-making, the more it is exposed to corruption pressures, and the more it distorts the very processes it was meant to monitor.

AI-based monitoring meets every condition for that effect. Keyboard and screen activity, presence at a desk, and the number of messages sent are all easy for a machine to count. When countable activity becomes a score and the score feeds into appraisals, employees face situations in which raising the score is faster than improving the work.

Figure 1 From activity score to target
AI scoresactivitykeys, desk,messagesUsed inappraisalsSignalsfakede.g. faketypingMetricdriftsWrongjudgmentsAI scores activitykeys, desk, messagesUsed in appraisalsSignals fakede.g. fake typingMetric driftsWrong judgments
Behaviour changes at the step where the score is used in appraisals and becomes a target. The more the changed behaviour satisfies the metric, the further the metric drifts from the quality of the work.

Behaviour aimed at the score is already on the record. In 2024 Wells Fargo dismissed more than a dozen employees after reviewing allegations of simulated keyboard activity that created the impression of active work. The wording comes from disclosures filed with FINRA, the securities industry's self-regulatory organization, and was reported by the trade publication AdvisorHub. Whether the employees used devices that move a mouse automatically is not known. What is known is that faking the activity record itself happened, and that it ended in dismissal.

03The evidence for performance gains is thin; the evidence for distrust and stress is not

Even if some employees fake activity, an employer could still accept that cost if monitoring raised performance overall. But when each expectation behind monitoring is checked against research, that premise does not hold up.

The central piece of evidence is a meta-analysis published in Personnel Psychology in 2023. A team of US researchers pooled 94 independent samples from studies of electronic performance monitoring. Its weight comes from the fact that it combines many studies rather than describing one workplace.

DimensionExpectation of those who monitorFindings of surveys and research
PerformanceMonitoring raises performancePooling 94 samples and 23,461 workers found no evidence of improvement (57% of monitored workers say it makes them more productive, a self-report)
TrustVisibility makes it easier to delegate51% of monitored workers feel they are not trusted
RulesMonitoring deters misconductMonitored employees broke more rules, such as taking unapproved breaks and ignoring instructions
BurdenNot consideredThe presence of monitoring is associated with higher stress regardless of its design (48% of monitored workers report stress)

The heaviest row is the one on rules. Monitoring is often introduced to curb misconduct, yet the study published in Harvard Business Review pointed in the opposite direction. The expectation that is easiest to cite as a reason for monitoring is the one the evidence contradicts most sharply.

This comparison matters now because monitoring is moving from recording activity to AI inference and scoring, and regulators have begun to name that step. The same meta-analysis found that organizations that monitor more transparently and less invasively can expect more positive attitudes from workers. Outcomes depend less on whether an employer monitors than on the design of the monitoring.

04The EU banned emotion inference at work and classed performance monitoring as high-risk

AI monitoring, as used in this column, is more than a system that records activity. It infers performance from those records, converts it into scores or rankings, and uses them in appraisals or in allocating tasks. Separating the recording step from the inference step shows where regulators have drawn their lines.

The European Union drew the clearest line. Its AI Act, adopted in 2024, prohibits AI systems that infer the emotions of a natural person in the workplace and in education, except for medical or safety reasons. The ban has applied since February 2, 2025. Violations can draw fines of up to 35 million euros or, for a company, 7% of total worldwide annual turnover, whichever is higher.

Monitoring that stops short of inferring emotions is regulated rather than banned. The AI Act classifies as high-risk any AI system used to monitor and evaluate the performance and behaviour of workers, or to allocate tasks based on individual behaviour or personal traits. Before putting such a system into service at the workplace, an employer must inform workers' representatives and the affected workers that they will be subject to it. That obligation was originally set to apply from August 2, 2026. An amending regulation published in the Official Journal in July 2026 moved it to December 2, 2027.

1

Japan

The Personal Information Protection Commission sets out points to observe when monitoring employees who handle personal data, in the Q&A to its guidelines. It says employers should ideally notify labour unions and consult them as needed when setting key rules.

2

New York State

Since May 7, 2022, employers that conduct electronic monitoring must give written notice upon hiring, obtain the employee's acknowledgment, and post the notice in a conspicuous place. Civil penalties start at 500 dollars for a first offense.

All three jurisdictions share one feature. Before specifying in detail what may be measured, they put in place a line that must not be crossed and a procedure for telling the people affected. They are building, through regulation, a way for workers to learn what the AI is turning into a score.

05Promotional review already monitors appropriateness, not the volume of activity

For people who create, review and use promotional materials in the pharmaceutical industry, monitoring is not a new word. Japan's Ministry of Health, Labour and Welfare issued guidelines in 2018 on the provision of sales information for prescription drugs. They require companies to set up a department that monitors the appropriateness of materials and of information-provision activities, independent of the departments that carry out those activities. Materials must be reviewed by this supervisory department before use and approved with the advice of a review and supervision committee.

DimensionActivity monitoringMonitoring under the guidelines
ObjectDevice input, presence and message logsAppropriateness of materials and activities
ObserverThe installed system and its administratorsA supervisory department independent of the business units
Link to appraisalAmount of activityWhether staff acted appropriately, or had others act appropriately

The two share a name but differ in purpose. What the guidelines feed into the evaluation of officers and employees is whether they acted appropriately, not how much activity they generated. If that distinction disappears when AI monitoring enters the review function, review work ends up being measured by volume.

Measuring by volume leaves out the core of review work. A reviewer reads the claims in a piece of material and the papers cited as evidence, then judges where they depart from the approved scope. Most of that does not register as activity a machine can count.

1

Time spent reading evidence

While a reviewer reads a paper, there is almost no device input. Counted by activity, the most important hours show up as blanks.

2

A decision to send back

A single item that took half a day of checking evidence still counts as one item.

3

What was missed

A serious miss appears neither in the number of comments nor in processing speed.

The concern is not whether an AI wrongly labels a reviewer as idle. In a review department scored on activity, fast, high-volume handling is rewarded, and slow work that reduces misses is not. That runs opposite to the role the guidelines assign to the supervisory department. The AI literacy worth building here is a habit: whenever you see a score produced by an AI, ask which signals it was built from.

06The harder quality is to see, the more visible activity gets scored in its place

Activity gets scored because the people doing the evaluating cannot see the quality of the work directly. In Microsoft's 2022 survey of 20,006 people in 11 countries, 85% of leaders said the shift to hybrid work had made it challenging to have confidence that employees were being productive, while 87% of employees said they were productive at work. The same report noted that hours worked, meeting counts and other activity metrics had in fact increased, and that some organizations track activity rather than impact.

When the evaluator cannot see quality, the quickest way to show results is to produce more visible activity. The evaluator judges by visible activity, and the evaluated produce more of it. As Campbell wrote, the more an indicator is used for decisions about people, the more it distorts the process it was meant to measure.

Figure 2 One design, three side effects
Appraise byactivityFaked activitydismissals over simulatedtypingMore rule-breakingHBR 2022 studyDistrust and stressAPA 2023 surveyAppraise by activityFaked activitydismissals over simulated typingMore rule-breakingHBR 2022 studyDistrust and stressAPA 2023 survey
All three come from the design choice of tying activity counts to appraisals, not from the accuracy of the monitoring.

The finding that monitored employees broke more rules and the survey result that half of monitored workers feel distrusted come from separate studies, but they point in the same direction. Employees receive monitoring as a statement of distrust, and that reception changes what they do next. Faked activity, rule-breaking and stress are not problems that better monitoring accuracy would solve. All three arise at once from a single design choice: tying activity counts to appraisals.

07Write down the purpose first, disclose it, and never count reading time as idle

The sequence an organization should follow when it introduces monitoring comes from the points in the Japanese commission's Q&A, plus one step a review function has to add for itself. Together they form the figure below.

Figure 3 Steps before and after monitoring
Specify thepurposeWrite rules,tell staffPick qualitymetricskeep readingtime inAudit thepracticeReport backto staffSpecify the purposeWrite rules, tell staffPick quality metricskeep reading time inAudit the practiceReport back to staff
Purpose, rules and audit are the points in the Japanese regulator's Q&A. Choosing quality metrics is not on that list; a review function has to add it.

The added step is choosing metrics of quality. For a review department, candidates include whether comments turned out to be valid, whether misses surfaced after approval, and whether issues sent back reappeared in the next piece of material. Any metric that counts reading time as idle is struck from the list. AI scores stay as input for people checking those metrics and are never used alone to decide appraisals.

Employees have decisions to make too. Leaving your work in a form the evaluator can read, for example recording each review decision together with its grounds, makes quality visible. It is different from faking. Manufacturing the activity signal itself, on the other hand, can be treated the same way as falsifying a record. The Wells Fargo case showed where that leads.

If the score does not match the work, the response is not to fake the signal but to ask, through the rules and the person responsible, what is being measured and for what purpose. I think being able to raise that question at work is a more useful practical ability in the age of AI than refining one's gamesmanship. The regulations described above have already begun to require employers to disclose that information.

Key Points ── 3 to take away
  1. AI monitoring measures activity easily and the quality of work poorly. When activity becomes the target, people shape how they look to the system and the metric stops reflecting quality. A meta-analysis pooling 94 samples found no evidence that monitoring improves performance.
  2. Regulation is forming around prohibitions and disclosure. The EU bans emotion inference at work and classes AI used to monitor and evaluate performance as high-risk, with those obligations now applying from December 2, 2027. In Japan, the data protection commission lists points to observe, including specifying and disclosing the purpose.
  3. Monitoring under Japan's sales-information guidelines looks at appropriateness, not volume. Protecting review quality means refusing metrics that count reading as idle time and judging outcomes. For employees, it means questioning the measure instead of faking the signal.
Closing

What AI scores is, in most cases, the appearance of work rather than its quality. Gamesmanship aimed at that appearance is a response to monitoring designed without any way to see quality, and the answer is not more skill at faking. It is deciding in advance what to measure and why, and telling the people who will be measured. For a promotional review department, that is the condition for keeping review quality safe from the score.

Sources & references
  1. The Wall Street Journal (Japanese edition). AIによる業務監視、従業員が出し抜くには (Callum Borchers). 2026-09-02. (Starting point of this column. Paywalled; only its premise is summarized.)
  2. American Psychological Association / The Harris Poll. 2023 Work in America Survey: topline data. 2023. (2,515 employed US adults; share monitored and how monitored workers view it.)
  3. Ravid, D. M., White, J. C., Tomczak, D. L., Miles, A. F., & Behrend, T. S. A meta-analysis of the effects of electronic performance monitoring on work outcomes. Personnel Psychology, 76, 5–40. 2023. (Meta-analysis of 94 samples, N = 23,461.)
  4. Thiel, C., Bonner, J. M., Bush, J., Welsh, D., & Garud, N. Monitoring Employees Makes Them More Likely to Break Rules. Harvard Business Review. 2022-06-27. (Rule-breaking by monitored employees; content confirmed via Arizona State University's summary.)
  5. Mattson, C., Bushardt, R. L., & Artino, A. R. When a Measure Becomes a Target, It Ceases to be a Good Measure. Journal of Graduate Medical Education, 13(1), 2–5. 2021. (Goodhart's original wording, Strathern's generalization, Campbell's formulation.)
  6. Hyman Cotter PC (Chicago Securities Law Blog). Wells Fargo Fires Workers for Allegedly Simulating Keyboard Activity to Fake Work. 2024-06. (Quotes AdvisorHub and the FINRA disclosure wording; use of devices unconfirmed.)
  7. Microsoft WorkLab. Hybrid Work Is Just Work. Are We Doing It Wrong? (Work Trend Index Special Report). 2022-09. (20,006 people in 11 countries; the gap between leaders and employees.)
  8. European Union. Regulation (EU) 2024/1689 (AI Act), Art. 5(1)(f), Art. 26(7), Art. 99(3), Annex III point 4; and amending Regulation (EU) 2026/1744 (Digital Omnibus on AI). Official Journal 2024-07-12 / 2026-07-24. (Ban on workplace emotion inference, high-risk classification and notice duty, postponed application.)
  9. Personal Information Protection Commission, Japan. Q&A on the Guidelines for the Act on the Protection of Personal Information (Q5-7). Accessed 2026-09-12. (Points to observe when monitoring employees.)
  10. Ministry of Health, Labour and Welfare, Japan. Guidelines on the Provision of Sales Information for Prescription Drugs (PSEHB Notification 0925-1). 2018-09-25. (Monitoring and pre-use review by an independent supervisory department; link to appraisal.)
  11. New York State Senate. New York Civil Rights Law § 52-c (Employers engaged in electronic monitoring). In force 2022-05-07. (Written notice upon hiring, acknowledgment, posting.)