Knowledgeworkspace
Compare AI agent work environments with a real file assignment
Choose through a practical file test: accurate outputs, usable revisions, clear approval boundaries and a handover your team can continue working with.
When you choose a working environment for AI agents, you should come away knowing which work your business can usefully delegate. A convincing demonstration answers only part of that question. What matters is whether your material becomes a correct, usable file, how you introduce a change and how much work remains for your team before adoption. A small test assignment makes these points tangible before you base a larger project on the environment.
You do not need a supplier ranking for this. You need your own standard, matched to the result you want. This article develops an acceptance test for preparing a small internal service catalogue. The dataset, conditions and example findings are entirely fictional. You can adapt the method without treating it as a performance promise for any particular solution. The test assesses a specific use case; it cannot establish suitability for every task across the company.
Describe the work that should be complete after the answer
Start with the handover to your team. If you need a cleaned overview, a good explanation of how to create one does not complete the agreed assignment. The expected file must actually exist, open correctly and contain the right information. A conversation can be very helpful along the way, for example when clarifying an incomplete source. The decisive question remains which implementation work someone still has to perform afterwards.
Before testing, write a sentence such as: “Our internal service overview should be created from the supplied table and be maintainable by the team leader afterwards.” That implies an editable format and a clear distinction between the input and revised output. If you only want alternative wording for a heading, you need a different test. The scale of the working environment does not, on its own, make it more appropriate for your purpose.
Identify the quality your business can assess. In the example, the team leader knows the permitted service names and can recognise missing responsibilities. They check the content. An office colleague opens the result in the spreadsheet application actually used by the team. You therefore evaluate both accuracy and usability. Looking only at the answer in the browser would not fully represent the subsequent work with the file.
Explore workspace for agent assignments that produce files and request access for your own test.
Determine the files and tools from the required result
Specify the input you will provide and the output you want back. Our example uses a small table as the source. The requested results are a cleaned table in an agreed editable format and a short reading version containing the main review notes. Establish whether the environment provides the necessary tools for this use. A long feature list does not replace an answer to that particular question.
Look for avoidable preparation work. Must you convert the data yourself first? Can the responsible person edit the output later? Does the process require applications your team does not use? These details may matter more than a particularly fast first response. Record the actions outside the session as well. Otherwise, you cannot see which work you really delegate and which work remains with your business.
Use artificial data for the first test. It should contain the typical difficulties of your case without requiring unnecessary real customer information. A perfectly clean dataset reveals little about handling duplicates or missing values. Conversely, an enormous collection of unrelated exceptions is not a manageable starting point. A few deliberate differences give you a checkable expectation for the later review, while keeping the task small enough to understand completely.
Use the same complete file assignment
The fictional dataset contains six rows with identifiers S01, S02, S03, S02, S04 and S05. Both S02 rows are identical in every field. S01 is “Maintenance”, S02 “Repair”, S03 “Commissioning”, S04 “Inspection” and S05 “Legacy service”. S01 through S04 are active; S05 is inactive. S03 has no responsible team entered. The other responsibility fields are complete in the test data. Identifiers and status are the business basis; similar words must not lead to additional invented services.
The assignment reads: “Create a cleaned internal service catalogue from the supplied table. Merge only duplicates that are identical in every field. Retain active and inactive services and show their status. Keep missing responsibilities visibly unresolved rather than inventing a person or department. Deliver the cleaned editable table and a short review note stating the number of input rows, unique services, active services and unresolved responsibilities.”
Add the boundary: “Do not change the input file. Create outputs in the agreed project area. Do not transfer anything into our actual internal library or send messages. Present the first output for review. Later adoption will be decided separately.” You are testing production of a concrete result. Permission to create a file in a test area is not permission to change existing company data or distribute the result elsewhere.
The expected calculation is simple and fully checkable. Six input rows contain one additional identical S02 row. Merging that duplicate leaves five unique services. Four are active and one is inactive. Exactly one service, S03, has an unresolved responsibility. The output must therefore contain five service rows and must not lose S05. These expectations apply only to the artificial dataset; they say nothing about the quality of your company’s real records.
Check the output before the agent’s explanation
Open the actual delivered table first. Find the five identifiers and check status and responsibility. The deliberately introduced cases are especially informative: does S02 appear once, is inactive S05 retained, and does S03 remain visibly unresolved? Then read the review note. A correct summary can accompany an incorrect file, so the explanation must not replace inspection of the output itself.
Suppose a fictional output reports five unique services but contains only four rows because the inactive service was removed. The test has not passed at that point. Retaining both active and inactive services was an explicit requirement. You can describe the error precisely without making a sweeping judgement about the entire environment. Equally, a brief explanation is sufficient when the file and review notes are correct and your team can work with them.
Check that identifiers remain intact during further handling. The team leader needs to find and update a service unambiguously. A polished reading version is insufficient for that purpose if the editable dataset is missing. Open both delivered versions and compare what they say. If the source had to remain unchanged, check that requirement against the original input or a comparison copy saved before the test began.
Test a revision and an agreed approval point
After the first correct output, give the same change in each test: “Sort active services before the inactive service. Within both groups, keep identifiers in ascending order. Add a note to the reading version explaining that responsibility for S03 must be clarified before internal use.” This feedback changes presentation and guidance rather than source values. Afterwards, check that all five services and their statuses remain. A revision must not quietly damage something that was already correct.
Observe whether you can redirect the work clearly. Is the change associated with the correct assignment? Can you identify which file version now applies? Do you have to explain the entire context again? These observations describe the workflow relevant to your business. Record the specific moment instead of simply writing “Good usability”. A visible activity history helps the decision only when your team can understand the current state and what needs to happen next.
For the approval point, the original boundary remains: the file may be prepared but must not yet enter the real company collection. Ask to see how that transition is presented for a decision in the agreed test process, without carrying it out. The review concerns a genuine boundary in your project. There is no need to trigger customer messages, publication or other consequential actions merely to demonstrate that a decision point exists.
Compare the complete effort through to adoption
Record your team’s actual time: preparing the source, arranging the permitted working area, explaining the assignment, checking, answering questions and adopting the result. Agent runtime is only part of the process. A fast output requiring extensive manual repair can create more total effort than expected. Conversely, a necessary question may prevent rework when it resolves an important uncertainty early enough to shape the result correctly.
Apply the same approach to costs. Use the actual terms offered for the relevant access and identify which support is included. Without specific information, do not assume free setup or fixed processing times. If the first attempt receives personal assistance, record that. A supported test does not automatically establish that your team can later perform the same process equally well without that assistance or under different access arrangements.
Avoid an arbitrary overall score that lets pleasant secondary features compensate for a missing essential requirement. A file that cannot be edited remains unsuitable for ongoing maintenance, even if the first response arrived quickly. Compare convenience and effort after the necessary conditions are met. This produces a useful business decision rather than a ranking assembled from many weakly justified points.
Clarify access, support and the actual handover
Before planning a project, establish how you can use the environment. Ask about the prerequisites for your specific assignment, the responsible contact and the route for questions. Identify who in your team needs access. A successful demonstration by someone else does not mean that the colleague who will take over the files already has the same access or can perform their part of the work.
The webRichtung workspace module page describes browser-based agent sessions with a live terminal, direction through agent input and completed files for download. Agents work in environments provided for the task, with approvals at critical points. Workspace is enabled on request. Use the test assignment to discuss the environment and handover intended for your use. The service catalogue exercise here is your own assessment method, rather than a promised dedicated import feature.
Have the handover checked by the person who will continue the work. They download the output, open it and find the unresolved responsibility. If they need to correct an identifier or add a note, the agreed editable format must support that work in their application. The test therefore finishes where the benefit begins: the work has reached the business and can continue without the original operator explaining every step.
Make a bounded decision and plan the next use
The outcome should support a clear statement: “This process is usable for preparing internal tables of this kind under these conditions.” Add the conditions, such as the available environment, agreed output format and content review by the team leader. Include any unresolved limitation. Do not generalise a successful small file assignment to untested system changes or work based on entirely different data.
If an essential requirement fails, identify the smallest useful next step. A presentation defect can be corrected within the same assignment and checked specifically. If the necessary environment is unavailable, clarify that prerequisite first. Do not repeat the whole comparison merely because a subjective wording choice could be different. The test should enable a commercial decision and end once its necessary requirements have been assessed with clear evidence.
For workspace, request access with that concrete need: the file you can provide, the result your team requires and the way you intend to check it. The module page provides the existing starting route. Establish access on request and begin with a bounded assignment whose result you can judge yourself. This turns an abstract platform question into a decision about usable work for your business.
Request workspace access and discuss your first file assignment and its handover.
Frequently asked questions
Do I need real customer data for the selection test?
No. A small artificial dataset with the same typical difficulties is sufficient initially. Suitability for other approved tasks should be assessed separately later.
Does a good conversational answer mean the test has passed?
When you commissioned a file, an actual usable file is part of the test. An explanation may accompany the result, but it cannot replace delivery.
How do I compare different levels of support?
Record which preparation and assistance were included and what your team had to do. Assess the complete route to a usable file and do not assume the same result under different access conditions.
How can I include workspace in my selection?
Describe your file assignment on the module page and request access. Workspace is enabled on request; discuss the appropriate working environment and test arrangements at that stage.