Future of Work

Can 70 Years of I-O Psychology Make Agents More Productive?

By Hugo Smith · August 20, 2026 · 7 min read

On July 31, 2025, DataRobot and NVIDIA announced the Agent Workforce Platform. The press release described a system that “lets organizations manage agents like digital employees, from deployment and integration to real-time oversight, retraining, and decommissioning.” The homepage calls it “the only end-to-end agent workforce platform for secure, scalable, production-grade agents” and says it is “built to power each phase of your agent lifecycle.”

The vocabulary is personnel vocabulary. Deployment, oversight, retraining, decommissioning. Digital employees. Agent lifecycle. isolved published an article listed as “AI Agents for HR: Hire Them Like You Would Anyone Else,” which argues that the CHRO should own the framework for bringing agents into an organisation in the same way HR owns the framework for hiring people. It applies hiring concepts directly: a job description for the agent, an evaluation of what it can do, and reference checks on where it has been deployed before.

This is a striking adoption, because the words carry a methodology behind them that the products do not implement. A job description in personnel psychology is the output of a job analysis: a structured inventory of tasks, linked to the knowledge and skills those tasks require, validated against a criterion measure of performance. An onboarding programme is built from that same task list. A performance review measures performance against a standard that was defined before the hire. The vocabulary is the visible surface of a method. Strip the method out and you are left with words that sound rigorous and carry none of the rigour.

What an agent job description actually is

In the vendor literature and in most build guides, an agent job description is a configuration document. It specifies the model, the system prompt, the tools the agent can call, the data it can reach and the guardrails around its behaviour. DataRobot’s platform organises these as templates with evaluation tools, task-specific benchmarks and policy controls, which is a genuine engineering accomplishment. It is a configuration spec, and calling it a job description borrows a label without importing the method that gave the label its meaning.

A job description produced by job analysis answers three questions that a configuration document does not: what tasks the role performs, what knowledge and skills those tasks require, and what measure of performance tells you whether someone in the role is doing them well. I have written about the first question elsewhere, and it remains the bottleneck on most agent projects. This piece is about the other two, and about a gap in the emerging agent-management literature that industrial-organisational psychology has been equipped to fill for decades.

Psychometrics on language models

Academics have begun running real psychometric instruments on language models. Guangyuan Jiang and colleagues at Peking University and UCLA developed the Machine Personality Inventory, adapting items from the IPIP-NEO Big Five instrument to test pre-trained language models. Their paper, “Evaluating and Inducing Personality in Pre-trained Language Models” (arXiv 2206.07550, submitted May 2022, accepted at NeurIPS 2023), showed that models produce different responses to personality items depending on how they were prompted, and that those differences are measurable.

The authors raise a caution that I-O psychologists will recognise. Model scores on personality items may reflect surface pattern-matching rather than a stable underlying construct. A model that scores high on conscientiousness items might be reproducing the statistical distribution of its training data rather than exhibiting anything that functions like a trait. This is the construct-validity argument, and it has been central to I-O psychology since the 1960s. The question is whether your measurement instrument captures the thing you think it captures or something else that happens to correlate with it.

I have written elsewhere about why personality measurement predicts less than demonstrations of actual work, in both people and models. The point here is different. Even researchers who are sympathetic to measuring model personality are the first to flag the construct-validity problem, and they flag it in the same terms I-O psychology has used for sixty years.

The task inventory that almost exists

The closest thing to a real task inventory of agent-suitable work came from computer scientists, not from I-O psychologists. Stanford’s SALT Lab published WORKBank, “Future of Work with AI Agents” (arXiv 2506.06576), a study of 1,500 domain workers across 104 occupations and more than 844 occupational tasks, built on the U.S. Department of Labor O*NET database, combining what workers want automated with capability assessments from AI experts. Erik Brynjolfsson and Diyi Yang are among the seven authors.

The study is a genuine empirical accomplishment. It identifies what work exists, breaks it into tasks, and asks whether each task is within reach of current AI. That is the front half of a job analysis. The half that is missing is the link from tasks to the knowledge, skills and abilities a role requires, and the criterion measure that tells you whether the agent is performing those tasks to standard. Those are the parts I-O psychology adds, and they are the parts that turn a task list into something you can hire, train or evaluate against.

Utility analysis and the person-versus-agent decision

There is a framework in I-O psychology for pricing a staffing decision. The Brogden-Cronbach-Gleser utility model takes a validity coefficient, a performance metric in dollar terms, the cost of the selection procedure and the number of hires, and produces an estimate of what the selection method is worth to the organisation. It has been used for decades to compare selection methods and to justify investment in better ones.

Organisations are now making a staffing decision that this framework was built for: whether a given task should be done by a person or by an agent. I searched for the Brogden-Cronbach-Gleser framework applied to the person-versus-agent staffing decision and did not find it. That search covered Google Scholar, the SIOP proceedings index and the usual HR-technology analyst outlets. The framework exists, the decision exists, and the bridge between them has not been built. People are being hired into roles whose purpose is to make this call, and the analytical tool most suited to pricing it has not been picked up.

What I-O psychology would add

If the method came with the vocabulary, three things would follow.

A task inventory would drive the agent specification. Instead of starting with a persona prompt or a model choice, the specification would begin with a structured list of what the agent needs to do, derived from observation of the work. The WORKBank study provides raw material for this. The I-O toolkit provides the method for turning it into a specification a builder can work to.

A criterion measure would sit behind the performance review. The dashboards that vendors call agent performance reviews monitor latency, cost, error rates and policy violations. Those are operational metrics. A criterion measure in the I-O sense is a measure of the quality of the work output, validated against what good performance looks like, and it is written before the agent is evaluated. The difference matters because a system can hit its operational targets and still produce work that is wrong, as anyone who has watched a model pass its benchmarks and fail in production has seen.

A utility analysis would price the make-or-buy decision. When the question is whether to deploy an agent or hire a person for a set of tasks, the Brogden-Cronbach-Gleser framework can turn that comparison into a number. The inputs are the validity of your selection method, the dollar value of performance in the role, and the cost of each option. The output is a decision that has been priced rather than argued by feel.

Where this leaves the conversation

The HR-technology market has adopted the language of personnel management for agents faster than it has adopted the method. DataRobot’s lifecycle language is real engineering, and isolved’s framing of agents as hires is a useful mental model for HR leaders. Both would be stronger if the vocabulary came with the technique behind it.

I-O psychology has been doing this work for seventy years. The task inventory, the criterion measure and the utility analysis are documented, validated and sitting in textbooks. The question is whether the people building agent platforms and the people buying them will reach for the method or settle for the words.


Sources: DataRobot, “DataRobot Announces Agent Workforce Platform Built with NVIDIA,” press release, July 31, 2025, datarobot.com/newsroom/press/datarobot-announces-agent-workforce-platform-built-with-nvidia/. · isolved, “AI Agents for HR: Hire Them Like You Would Anyone Else” (blog listing title), published as “The Hire Nobody Made,” August 5, 2026, isolvedhcm.com/blog/ai-agents-for-hr-hiring-process. · Jiang, G., Xu, M., Zhu, S.-C., Han, W., Zhang, C., Zhu, Y. (2022/2023). Evaluating and Inducing Personality in Pre-trained Language Models. arXiv 2206.07550. Accepted at NeurIPS 2023. · WORKBank: “Future of Work with AI Agents,” arXiv 2506.06576. Authors include E. Brynjolfsson and D. Yang (Stanford SALT Lab). Figures per the abstract: 1,500 domain workers, 104 occupations, over 844 tasks, built on ONET.*

Disclosure: some links above are partner links. If you sign up through them we may earn a commission, at no cost to you. It never changes our verdict; see our disclosure policy.