Text me
All posts

I Built a Safety Layer for AI-Generated Bash

A verifier, a sandbox, and a refusal default. The story behind safe-cli and why AI agents need a gate between them and the terminal.

View safe-cli on GitHub

There is something slightly uncomfortable about giving an AI agent access to a terminal.

On one hand, it is incredibly useful. An AI can inspect a project, install dependencies, modify files, run tests, move things around, and automate tasks that would normally take a developer quite a bit of time.

On the other hand, the terminal does not really care whether the person, or model, typing the command understands what it is doing.

A command can be syntactically valid, look completely reasonable, and still do something you did not intend.

That was the idea behind safe-cli.

I wanted a way to put a verification layer between AI-generated Bash and the actual machine.

Not another prompt asking, "Are you sure?"

Something that could actually look at the command, test it, and refuse to run it when it crosses a safety boundary.

The problem with trusting generated shell commands

AI coding agents generate shell commands constantly.

Most of those commands are completely fine.

The problem is the small percentage that is not.

Take something as simple as:

rm -rf $TARGET

At a glance, it looks like a normal command.

But if $TARGET is not what you think it is, you are potentially deleting something completely different.

Or consider:

eval "$user_input"

That might look like a convenient way to execute something dynamically, but it also creates an obvious command-injection problem.

Then there are the less dramatic failures.

An unquoted variable.

A glob that expands differently than expected.

A missing error check.

A script that enters an infinite loop.

A command that creates or deletes files somewhere it should not.

These are not necessarily things an AI "does not know."

They are things that can happen when generating code at scale, where the output is often accepted and executed much faster than a human would manually review every line.

That creates a different problem than traditional software development.

The question is not just whether the code is correct. The question is whether the code should be trusted enough to execute.

That is where safe-cli comes in.

The basic idea is intentionally simple

AI-generated Bash
       ↓
   safe-cli
       ↓
    Verify
       ↓
 ┌─────┴─────┐
 ↓           ↓
Pass        Fail
 ↓           ↓
Execute     Refuse

Instead of letting a script go directly from an AI agent to the operating system, safe-cli acts as a gate.

The command has to make it through verification first.

If verification fails, execution stops.

There is not a second path around the verifier.

There is not a "warning, but continue anyway" path for blocking failures.

The goal is to make the dangerous mistake fail before it reaches the machine.

I did not want one safety check

One of the things I found interesting while building this was that there is not really one perfect way to decide whether a Bash script is safe.

A syntax checker can tell you whether the shell understands the script.

It cannot necessarily tell you whether the script does what you intended.

A static analyzer can find suspicious patterns.

It does not actually execute the code.

A sandbox can show you what happens when the script runs.

But a sandbox by itself does not necessarily tell you that the script contains a dangerous construction.

So instead of trying to find one magical analyzer, safe-cli uses several different layers.

It currently runs Bash through ten independent verification layers, including Tree-sitter parsing, native Bash syntax checking, ShellCheck, formatting analysis, language-server diagnostics, behavioral tests, Docker sandbox execution, adversarial inputs, fuzzing, and filesystem side-effect detection.

The thinking is pretty straightforward:

If one perspective misses something, another perspective might catch it.

And if a layer reports a hard failure, the command does not get executed.

Testing the code is not enough

This is probably the part of the project I care about most.

There is a difference between asking:

"Does this script look correct?"

and asking:

"What actually happens when I run this?"

Those are two different questions.

That is why safe-cli does not stop at static analysis.

The runtime portion can execute the script inside a Docker sandbox with restrictions such as no network access, a read-only filesystem, a non-root user, dropped capabilities, and a timeout.

The idea is to make the execution environment itself hostile to bad behavior.

If something gets stuck, it gets killed.

If something tries to depend on network access, there is not supposed to be a network available.

If something tries to make unexpected filesystem changes, the side-effect layer can detect them.

This turns verification from something purely theoretical into something closer to an experiment:

"Okay, let us actually run it, but somewhere controlled first."

Then there is the adversarial side

Normal tests tend to ask whether software works with the inputs we expect.

Security testing asks what happens when we give it inputs we do not expect.

safe-cli has specific adversarial and fuzzing layers for this reason.

For example, it can take hostile characters and inputs and try to expose command-injection problems.

The project currently includes dozens of adversarial quoting inputs as well as property-based fuzzing around shell metacharacters.

That matters because shell scripting has an uncomfortable relationship with strings.

A string is not always just a string.

Depending on how it is handled, it can become:

  • a filename,
  • multiple arguments,
  • a glob,
  • a command,
  • or something that gets interpreted again.

A lot of shell security problems come from crossing that boundary accidentally.

So instead of simply trusting that a function handles its input correctly, safe-cli tries to attack the assumption.

Refusing is part of the design

A big part of this project is the decision to make refusal a first-class behavior.

Most developer tools are designed to help you continue.

A compiler tells you what is wrong so you can fix it.

A linter gives you warnings.

A formatter changes the code.

A test suite tells you what failed.

safe-cli has a slightly different responsibility.

Its job is sometimes to say:

No. Do not run this.

That sounds obvious, but it changes the architecture.

A verification failure is not just information.

It is a control decision.

The documented exit behavior makes that explicit: a verification failure returns a non-zero status and the script is not executed.

That is especially important when the caller is an AI agent.

An AI agent can see an error and decide to try something else.

That is actually useful.

What I do not want is for the agent to see a safety failure and simply continue with the dangerous command anyway.

I also wanted the tool to be practical

Security tools have a tendency to become annoying.

If using the secure version of something requires ten extra commands, complicated configuration, and a manual approval process every time, developers eventually stop using it.

So the interface is intentionally small.

You can verify a script:

safe-cli verify script.sh

You can verify and execute:

safe-cli run script.sh

You can provide Bash directly:

safe-cli exec 'echo hello'

And there is a repair workflow:

safe-cli fix script.sh

There is even a doctor command for checking whether the installation itself is healthy.

The goal is for the safety layer to become part of the normal workflow rather than something developers remember to use only when they are worried.

The repair system is deliberately conservative

One thing I did not want was a system that "fixes" code by changing its behavior until the tests happen to pass.

That is dangerous in its own way.

If an automated repair system can rewrite arbitrary code, it can theoretically make a script look safer while quietly removing functionality.

So the repair system is intentionally constrained.

It can apply minimal-difference fixes for certain warning-level problems, but it does not automatically repair syntax errors or informational diagnostics, and it does not delete functionality simply to make a test pass. Backups are created before edits.

In other words:

Fixing the code should not mean changing what the code is supposed to do.

The deeper idea

The more I worked on this, the more I started thinking about safe-cli as something slightly different from a Bash security tool.

It is really an experiment in trust boundaries for AI agents.

We are increasingly giving AI systems the ability to interact with the real world.

At first, that meant generating text.

Then code.

Then code that could be executed.

Then tools that could modify files, install software, interact with APIs, and operate computers.

That progression creates a problem:

How much should an agent be trusted simply because it produced something that looks reasonable?

safe-cli takes a very conservative answer:

It should not have to be trusted blindly.

The output should have to pass through another layer.

The agent can propose an action.

The verifier evaluates it.

The sandbox tests it.

The safety rules enforce boundaries.

And only then does the command get a chance to touch the real environment.

It is not about making AI perfect

I do not think the interesting goal is to make AI-generated Bash perfect.

That is probably impossible.

The more realistic goal is to make mistakes less catastrophic.

If an AI generates a slightly inefficient command, that is annoying.

If it generates a command that accidentally deletes the wrong directory, that is a different class of problem.

The purpose of a safety layer is to catch that difference.

It is essentially accepting the premise that AI will make mistakes, and designing the execution environment so that those mistakes have somewhere to stop.

Where I see safe-cli going

The project is still an experiment, but I think the underlying idea is bigger than Bash.

Bash just happens to be a particularly obvious place to start because shell commands have immediate access to the operating system.

The broader pattern is:

Generated action
       ↓
   Verification
       ↓
 Controlled execution
       ↓
   Observation
       ↓
     Result

That is a pattern I think we are going to see more of as AI agents become increasingly autonomous.

Instead of asking whether an agent is trustworthy, we can start asking a different question:

What happens when the agent is not?

That is the mindset behind safe-cli.

It is not trying to make the AI stop making mistakes.

It is trying to make those mistakes harder to turn into real damage.