Summary
On a macOS 26 host, output from tart exec in a macOS 12–15 guest breaks once more than a few hundred KiB is sent quickly. Depending on timing, one of three things happens:
tart exec fails with Error: unavailable (14): Transport became inactive.
tart exec hangs.
- The output is silently truncated and
tart exec still returns the guest command's real exit code, including 0.
The same host, Tart build, and guest agent binary work perfectly with a macOS 26 guest (1 GiB in ~3 s, checksum matches).
I took the guest agent and gRPC out of the path with a raw AF_VSOCK sender in the guest, and the data loss still happens. So I believe the loss is in the virtio-vsock transport between older guest kernels and Virtualization.framework, not in Tart or the agent. Tart can't fix that layer, but it could detect the loss instead of reporting success (see "Suggestions").
Environment
- Host: macOS 26.6.2 (25G83), Apple silicon; Tart 2.37.0, and also reproduced on
main at 4e58a2a
- Guests:
ghcr.io/cirruslabs/macos-{monterey,sonoma,tahoe}-vanilla plus tart-guest-agent 0.14.1-cb39b12, installed as the usual LaunchAgent (--run-agent) and LaunchDaemon (--run-daemon). The agent binary is identical on every guest.
| Guest |
Result |
| macOS 12.7.6 |
1 MiB: truncated to 330–1,008 KiB; about half the runs exit with the command's own code |
| macOS 14.8.7 |
1 MiB: fails after ~240–550 KiB |
| macOS 26.6.2 |
1, 16, 64 MiB, 1 GiB, 4 GiB: all intact |
An earlier run on the same host also failed on macOS 13.7.4 (1 MiB) and macOS 15.7.7 (16 MiB; 1 MiB passed). I didn't re-test those two for this report.
Steps to reproduce
tart run --no-graphics <sonoma-or-monterey-vm-with-agent> &
tart exec <vm> sh -c 'dd if=/dev/urandom of=/tmp/g.bin bs=1m count=4 2>/dev/null'
tart exec <vm> head -c 1048576 /tmp/g.bin | wc -c # expect 1048576
tart exec <vm> sh -c 'head -c 1048576 /tmp/g.bin; exit 7' > out.bin; echo $?; wc -c < out.bin
Monterey 12.7.6, six runs of the last command:
rc=1 bytes=258048 Error: unavailable (14): Transport became inactive
rc=7 bytes=843776
rc=7 bytes=876544
rc=7 bytes=929792
rc=1 bytes=724992 Error: unavailable (14): Transport became inactive
rc=1 bytes=757760 Error: unavailable (14): Transport became inactive
The rc=7 runs show that tart exec received the exit event after data went missing. Whole HTTP/2 DATA frames disappeared while the framing stayed in sync, so neither gRPC nor tart exec noticed. With exit 0, the same thing looks like success.
Isolation: the loss happens below the agent and gRPC
On Sonoma and Tahoe clones, I replaced the agent's port-8080 listener with a minimal Go AF_VSOCK server. It writes a known byte pattern (byte i = i mod 251) and logs how much the guest kernel accepted. On the host, I connected straight to the VM's control.sock and checked every byte. The two programs are below.
| Guest |
Size / write size / pacing |
Guest write() accepted |
Host received |
| 14.8.7 |
1 MiB / 64 KiB / none |
1,048,576 in 0.6 ms, no errors |
393,216, first wrong byte at 262,144 |
| 14.8.7 |
1 MiB / 4 KiB / none |
1,048,576 in 1.8 ms, no errors |
393,216 (clean truncation) |
| 14.8.7 |
1 MiB / 16 KiB / none |
n/a |
458,752, first wrong byte at 327,680 |
| 14.8.7 |
16 MiB / 1 MiB / none |
n/a |
524,288, first wrong byte at 262,144 |
| 14.8.7 |
1 MiB / 128 KiB / 5 ms sleep |
n/a |
all, intact |
| 14.8.7 |
4 MiB / 64 KiB / 20 ms sleep |
4,194,304 |
all, intact |
| 26.6.2 |
every case above |
n/a |
all, intact |
When the guest sends quickly, the host loses 64 KiB-aligned blocks after about 256 KiB, and the stream then stalls. The guest kernel accepted every byte without error. The host-side proxy (ControlSocket) is the same code that handles Tahoe at over 300 MiB/s. That points at virtio-vsock flow control between macOS ≤15 guest kernels and the macOS 26 host. I haven't tried an older host to see whether it's a regression on the host side.
Guest sender (Go; builds inside the tart-guest-agent module to reuse internal/vsock)
package main
import (
"bufio"
"fmt"
"io"
"os"
"time"
"github.com/cirruslabs/tart-guest-agent/internal/vsock"
)
// Reads "SIZE CHUNK SLEEPMS\n" and then writes SIZE bytes (byte i = i%251).
func main() {
l, err := vsock.Listen(8080)
if err != nil {
fmt.Fprintln(os.Stderr, "listen:", err)
os.Exit(1)
}
for {
c, err := l.Accept()
if err != nil {
continue
}
go func() {
defer c.Close()
r := bufio.NewReader(c)
var size, chunk, sleepMs int
if _, err := fmt.Fscanf(r, "%d %d %d\n", &size, &chunk, &sleepMs); err != nil {
return
}
go io.Copy(io.Discard, r)
buf := make([]byte, chunk)
sent := 0
for sent < size {
n := min(chunk, size-sent)
for i := 0; i < n; i++ {
buf[i] = byte((sent + i) % 251)
}
w, err := c.Write(buf[:n])
sent += w
if err != nil {
fmt.Fprintf(os.Stderr, "write error after %d: %v\n", sent, err)
return
}
time.Sleep(time.Duration(sleepMs) * time.Millisecond)
}
fmt.Fprintf(os.Stderr, "sent %d bytes\n", sent)
}()
}
}
Host reader (Python)
import os, socket, sys
vm, size, chunk, sleep_ms = sys.argv[1], *map(int, sys.argv[2:5])
s = socket.socket(socket.AF_UNIX); s.settimeout(6)
s.connect(os.path.expanduser(f"~/.tart/vms/{vm}/control.sock"))
s.sendall(f"{size} {chunk} {sleep_ms}\n".encode())
got, bad = 0, -1
try:
while (b := s.recv(1 << 20)):
if bad < 0:
bad = next((got + i for i, x in enumerate(b) if x != (got + i) % 251), -1)
got += len(b)
except TimeoutError:
pass
print(f"got={got}/{size} first_bad={bad}")
To run this, stop the guest's org.cirruslabs.tart-guest-agent LaunchAgent so the sender can bind port 8080.
Suggestions
- Make loss detectable.
tart exec currently has no way to notice missing stdout or stderr. If the agent's Exit event carried the total stdout and stderr byte counts it sent (or IOChunk carried an offset), tart exec could fail with a clear error instead of returning the command's exit code over truncated output. This would need matching changes in tart-guest-agent and Tart.
- Latent exit-0 path. In
Exec.execute(), if the response stream ends without an exit event, the task group finishes normally and tart exec exits 0 (Exec.swift#L188-L212). That isn't what caused the runs above, but a stream with no exit event should probably be an error.
- The vsock loss itself looks like an Apple issue. If you already know about it or have a Feedback number, I'm happy to add my results there. If a workaround exists (for example, a host-side setting), it would be good to document it for macOS ≤15 guests.
Possibly related: #1337 (clipboard fails on Sonoma guests but works on Tahoe guests with the same agent). That uses a different virtio device, but it has the same pattern of an older guest failing on a new host. The control socket also stops accepting connections after some of these failures; I've filed that separately as #1346.
Context
I found this through a tool that reads build logs out of verification VMs with tart exec … cat. On macOS 12–15 guests, those logs could come back truncated with no error.
This report was investigated, reproduced on my machine, and written up with the help of Claude (Anthropic's AI assistant). I reviewed it before posting.
Summary
On a macOS 26 host, output from
tart execin a macOS 12–15 guest breaks once more than a few hundred KiB is sent quickly. Depending on timing, one of three things happens:tart execfails withError: unavailable (14): Transport became inactive.tart exechangs.tart execstill returns the guest command's real exit code, including 0.The same host, Tart build, and guest agent binary work perfectly with a macOS 26 guest (1 GiB in ~3 s, checksum matches).
I took the guest agent and gRPC out of the path with a raw
AF_VSOCKsender in the guest, and the data loss still happens. So I believe the loss is in the virtio-vsock transport between older guest kernels and Virtualization.framework, not in Tart or the agent. Tart can't fix that layer, but it could detect the loss instead of reporting success (see "Suggestions").Environment
mainat 4e58a2aghcr.io/cirruslabs/macos-{monterey,sonoma,tahoe}-vanillaplus tart-guest-agent 0.14.1-cb39b12, installed as the usual LaunchAgent (--run-agent) and LaunchDaemon (--run-daemon). The agent binary is identical on every guest.An earlier run on the same host also failed on macOS 13.7.4 (1 MiB) and macOS 15.7.7 (16 MiB; 1 MiB passed). I didn't re-test those two for this report.
Steps to reproduce
Monterey 12.7.6, six runs of the last command:
The
rc=7runs show thattart execreceived the exit event after data went missing. Whole HTTP/2 DATA frames disappeared while the framing stayed in sync, so neither gRPC nortart execnoticed. Withexit 0, the same thing looks like success.Isolation: the loss happens below the agent and gRPC
On Sonoma and Tahoe clones, I replaced the agent's port-8080 listener with a minimal Go
AF_VSOCKserver. It writes a known byte pattern (byte i = i mod 251) and logs how much the guest kernel accepted. On the host, I connected straight to the VM'scontrol.sockand checked every byte. The two programs are below.write()acceptedWhen the guest sends quickly, the host loses 64 KiB-aligned blocks after about 256 KiB, and the stream then stalls. The guest kernel accepted every byte without error. The host-side proxy (
ControlSocket) is the same code that handles Tahoe at over 300 MiB/s. That points at virtio-vsock flow control between macOS ≤15 guest kernels and the macOS 26 host. I haven't tried an older host to see whether it's a regression on the host side.Guest sender (Go; builds inside the tart-guest-agent module to reuse
internal/vsock)Host reader (Python)
To run this, stop the guest's
org.cirruslabs.tart-guest-agentLaunchAgent so the sender can bind port 8080.Suggestions
tart execcurrently has no way to notice missing stdout or stderr. If the agent'sExitevent carried the total stdout and stderr byte counts it sent (orIOChunkcarried an offset),tart execcould fail with a clear error instead of returning the command's exit code over truncated output. This would need matching changes in tart-guest-agent and Tart.Exec.execute(), if the response stream ends without anexitevent, the task group finishes normally andtart execexits 0 (Exec.swift#L188-L212). That isn't what caused the runs above, but a stream with no exit event should probably be an error.Possibly related: #1337 (clipboard fails on Sonoma guests but works on Tahoe guests with the same agent). That uses a different virtio device, but it has the same pattern of an older guest failing on a new host. The control socket also stops accepting connections after some of these failures; I've filed that separately as #1346.
Context
I found this through a tool that reads build logs out of verification VMs with
tart exec … cat. On macOS 12–15 guests, those logs could come back truncated with no error.This report was investigated, reproduced on my machine, and written up with the help of Claude (Anthropic's AI assistant). I reviewed it before posting.