Skip to content

From Workflow to Trusted System: Going Deeper with AI

How in-house lawyers can responsibly scale AI workflows from sandbox testing to agentic systems.

Authors

  • Laura Belmont

    General Counsel

    The L Suite

Lloyd AI

In an earlier blog post, my message to in-house counsel was that you're not behind on using AI and the on-ramp is smaller than it may seem: pick a task, give it fifteen minutes, build from there. If you did that and set up a morning triage, a product-review skill or another workflow, you've felt the thing the early adopters are compounding. You've also probably bumped into the next set of questions that don't have a "just start" answer.

This worked on something low-stakes. When is it safe to point at real, confidential work? How much should I let an agent do? And how do I scale this for my team?

These are the right questions to be asking. Going deeper is about building systems you can trust with work that matters.

The Sandbox Is for Comfort...and Diligence

During the Claude for Legal: In-House Edition (Part 1) webinar, Heather Stevenson offered advice for the hesitant: open a personal account and test with non-sensitive information. As she put it, "I have a personal Claude account on my own computer with much less sensitive stuff that I played with almost everything I could first... you'll see that it's less scary than it may initially seem." That's a good starting move. And there's also a more strategic reason to work in a sandbox that matters as you go deeper.

Testing is where you learn a tool's limitations on data that doesn't matter, so you can make an informed risk/reward judgment when it comes to data that does. You run it against cases you already know the answer to, find where it's reliable and where it isn't, and only then decide whether the upside justifies the exposure.

It's worth being clear about what "sandbox" means, because there are different ways you can go about this. The first is working with dummy or public data in an ordinary account: no confidential or sensitive information, so a bad output costs nothing. It means answering the following questions before any confidential matter touches the tool:

  • Where does the tool get things wrong and how would I catch that in real work?

  • What does it do well enough that I'd actually rely on it?

  • What's the setup that makes it perform and why does that setup matter? Understanding the why is what lets you judge whether a given configuration is worth it.

Claude for Legal: In-House Edition (Part 1)

Watch an on-demand, practical session exploring how in-house legal teams are using Claude for Legal to get real work done.

Get the Recording

The second is a true technical sandbox environment: an isolated, secure space where software or code runs separated from your main operating system and network, so that any errors or malicious activity inside it can't reach the host system.

The first lets you safely judge the tool's output; the second lets you safely run the tool itself. As you move toward agents and connected systems, the distinction starts to matter.

This logic carries directly into the business context. Your company doesn't want to go through the work of onboarding a particular tool for a use case unless they already know it's worth it and can actually serve the business purpose. A sandbox (dummy data, an isolated environment or both) lets you find that out first.

Deciding When a Workflow Is Trustworthy Enough for Confidential Work

At some point you'll want to move from public and dummy data to the real thing. A few questions worth answering first, some of which the sandbox has already taught you:

  • Data handling. Where is your sensitive data stored, what is the underlying permissioning structure, and what enterprise controls are in place for the AI tool (e.g., do you have ZDR)? Has your organization approved this specific use of the tool — not just the tool in general?

  • Verifiability. Can you trace the output back to a source or is it the AI making assertions you can't check? This is where grounding matters. Shaun Sethna, CLO of The L Suite, made the point about benchmarking his contract playbooks against peer knowledge through the L Suite's Lloyd connector: "You can click here and it comes to a thread on the Braintrust that shows what the actual discussion was. It's not just making things up, it's taking you back to the Braintrust." Output you can trace is far safer to rely on than confident prose with no provenance.

  • Reversibility. If the workflow gets something wrong, how bad is it, and can you catch it before it matters? A summarization task you'll review is low-risk; an agent that sends external communications is not.

This is about making the confidential-data decision deliberately, with the same risk/reward lens you'd apply to any vendor or process.

When the Tool Stops Answering and Starts Acting

Mark Pike, Associate General Counsel at Anthropic, named the ceiling most people hit: "A lot of lawyers stop [with chat]. They just ask a question, get a response, and think that's the bulk of what AI can do." But with agentic AI, the tools take actions. Agents can read your inbox, draft and send, move files, and run a process end to end on a schedule.

Roman Perchyts, General Counsel at airSlate, described what that looks like in practice: he built a product-review agent that scans a Slack channel, runs a preliminary analysis, and delivers a pre-researched memo before he's seen the message. "By the time you touch it," he said, "there's some work that's already been done." That's the leverage. It's also where the risk profile changes completely, because a tool that can act can act wrongly.

This is where your lawyerly instincts and training are an advantage. You already think in terms of authority and scope, and that maps almost directly onto how much latitude to give an agent.

A useful distinction is read versus write. An agent with read access can analyze, summarize, triage, and surface (like Roman's memo-before-you-touch-it) without the power to change anything. An agent with write access can send, delete, edit, or commit on your behalf. Start with read-only. Grant write privileges narrowly and only once you've seen enough to trust it; and even then, scope them to specific actions rather than handing over full rein.

The sandbox and the agent question converge here: the place to learn how an agent behaves, and how to constrain it, is on a personal or low-stakes task where a mistake costs nothing. By the time an agent is doing something that matters, you've already learned where its guardrails need to be.

Where This Leaves You

The lawyers who'll be most valuable over the next few years aren't the ones who adopted AI the fastest. They're the ones who learned to deploy it with judgment, which is, conveniently, the thing you were already trained to do. Start in the sandbox. Keep agents on a short leash until they earn a longer one. Make the confidential-data call deliberately.



This post draws on a live practitioner session from The L Suite's Claude for Legal In-House Edition webinar (Part 1), featuring Mark Pike (Associate General Counsel, Anthropic), Heather Stevenson (General Counsel, Red Cell Partners), Roman Perchyts (General Counsel, airSlate), and Shaun Sethna (CLO, The L Suite). Part 1, "From Chat to Cowork: How In-House Lawyers Are Building AI Workflows," is here.