Collect archival data
Most of the output can be produced by software today.
An agent can find, read and summarise the sources this work depends on, and it can cover far more of them than a person would ever open. What it does not do is decide which sources deserve trust, or what the findings mean for your situation.
| How automatable | Mostly automatable |
|---|---|
| Kind of work | Research and retrieval work |
| Jobs that do it | 8 occupations |
| Tool categories that apply | Web, data and docs, Generalist agents |
| O*NET activity | Getting Information (4.A.1.a.1) |
What this looks like on the job
Not our paraphrase. These are real task statements recorded against this activity, spread across the 8 occupations that perform it.
- Track the flow of work forms, including in-house data flow or electronic forms transfer.
- Obtain and study medical, psychological, social, and family histories by interviewing individuals, couples, or families and by reviewing records.
- Consult site reports, existing artifacts, and topographic maps to identify archeological sites.
- Locate and obtain existing geographic information databases.
- Gather historical data from sources such as archives, court records, diaries, news files, and photographs, as well as from books, pamphlets, and periodicals.
- Read and study reports in order to compile information and data for geological and geophysical prospecting.
- Conduct internet-based and library research.
- Prepare archival records, such as document descriptions, to allow easy access to information.
- Collect detailed information on individuals for use in biographies.
- Interview individuals, and research public databases in order to obtain information.
What we would actually use
One recommendation rather than a shortlist, because a shortlist is just your problem handed back. This is what we would buy for research and retrieval work, and what it costs.
Research is the one shape of work that needs no integration at all, so a single $20 assistant with browsing is the whole stack. Add a crawler only once you are pulling hundreds of pages on a schedule rather than reading a few.
Chosen for the cheapest thing that does the job, not the best funded. Nothing on this site is sponsored, and the full landscape is there when you want to disagree with us.
The rest of the category
Context, not alternatives to weigh up. We name the layer rather than promise a named product does your specific task.
Generalist agents
Matched from the kind of work, not from vendor marketing. Check any of them can reach the system your records actually live in — that connection, not the model, is where these projects stall.
How people actually do it
No workflow names this activity yet, but these automate the same kind of work, with the prompts to paste and what to keep for yourself.
Who does this work
8 occupations in the O*NET database perform this activity. Each one has a full breakdown of its other tasks.
Work that goes with it
O*NET groups these under “Gather information from physical or electronic sources”. In practice they tend to be done by the same person, in the same sitting.
Method. The activity, its taxonomy placement and the occupations that perform it come straight from the public-domain O*NET 30.3 database. The automatability band is ours: we map each of O*NET’s 41 generalized work activities to how much of its output current software can produce, assuming a person still reviews and owns the result. It is a coarse three-way judgment on purpose. A precise-looking percentage here would be invented.