Actors are the unit of long-running services in Lambda Prelude. They are implemented on OCaml 5 effect handlers — no Eio, no Miou, no POSIX-only primitives. The same actor code runs on Windows and Linux, under the interpreter and under AOT.

Three Communication Primitives

The actor surface starts with three communication primitives — spawn, send, receive — plus Actor self, which returns the current fiber's actor ref.

Actor self                  "actor ref of the current fiber"
Actor spawn: [:me | body]   "spawn a new fiber, body sees its own ref"
actor ! message             "fire-and-forget send"
actor receive               "block until a message arrives"

A first ping-pong:

parent := Actor self.

child := Actor spawn: [:me |
  msg := me receive.
  msg printNl.
  parent ! #pong
].

child ! #ping.
parent receive printNl.

State Lives in Loop Arguments

There is no shared mutable state. An actor's state is the argument of its recursive loop. To "update" state, the actor calls itself with a new value:

Object subclass: #Counter.

Counter class >> run: me count: n =
  me receive: {
    Inc      -> [:msg | Counter run: me count: n + 1].
    GetCount -> [:msg | msg reply ! n. Counter run: me count: n]
  }.

counter := Actor spawn: [:me | Counter run: me count: 0].

n is closed inside the call frame; no other fiber can read or write it.

Typed Selective Receive

me receive: { Class -> [:msg | ...] } is a class-tag dispatch table. The scheduler scans the mailbox head-to-tail and runs the first matching handler. Non-matching messages stay in their original positions — selective receive.

Object subclass: #Inc.
Object subclass: #Get fields: (reply).

Counter class >> run: me state: n =
  me receive: {
    Inc -> [:msg | Counter run: me state: n + 1].
    Get -> [:msg | msg reply ! n. Counter run: me state: n]
  }.

The dispatch table threads each message class through Hindley-Milner inference, so the handler body sees concrete types — no Obj.magic, no untyped escape.

Receive Timeouts

receive:after:do: returns either a matched message or the timeout branch.

me receive: { Tick -> [:t | handle: t] }
   after: 1000
   do:    [Log warn: 'no tick for 1 s'].

Futures and ask:

ask: is request / response sugar on top of send + receive. It returns a Future that resolves when the reply arrives.

fut := worker ask: (Compute fields: { x: 7 }).
fut await printNl.

Manual Future construction is also available; ask: is the common path.

Failure semantics follow the Erlang vocabulary, applied to the local model:

Object subclass: #Boom.
Object subclass: #Worker.

Worker class >> loop: me =
  me receive: {
    Boom -> [:msg | Error raise: 'kaboom!']
  }.

main   := Actor self.
worker := Actor spawn: [:me | Worker loop: me].

main monitor: worker.
worker ! (Boom fields: {}).

main receive: {
  Down -> [:d |
    'Down received' printNl.
    d reason printNl
  ]
}.

'main still alive' printNl.

Down carries the dead actor's ref and the failure reason. The observer keeps running — there is no cascade unless the actor is also linked.

Actor spawn: and the monitor: that follows are two steps, and on a worker domain the child can run and die in between. Registering on an actor that already died therefore reports the death instead of dropping the registration: monitor: answers with a Down carrying the reason recorded at death, and linkTo: raises — or delivers Exit, when the caller traps exits. A late observer sees what an early one would have.

An actor that dies abnormally with no monitor and no link is reported on stderr. Such a death used to leave no trace: the reason was recorded for observers, and with none it was dropped, so a spawned worker that raised simply stopped existing and the only evidence was work that never happened. --quiet turns the report off, together with the unmatched-message warning.

Supervision

Supervision trees are written as plain classes, not as dedicated syntax. A supervisor monitors its children and restarts them on Down. When the restart budget is exhausted, the supervisor stops; if it is itself monitored by a higher supervisor, that level takes over from there.

The Supervision module in stdlib provides reusable strategies (one-for-one, one-for-all, rest-for-one) and restart policies, but the underlying mechanism is just monitor + spawn.

Multi-Domain Scheduling

The runtime is multi-domain: actors can spawn across OS threads with cross-domain inbox handoff and work-stealing. The same runtime is compiled into the AOT binary, so it gets the identical multi-domain scheduler — CPU-bound workloads spread across cores in both modes, with no separate single-domain runtime.

Bounded Socket I/O

TCP connect:port:timeout:, TCPSocket >> readLine:timeout: / readBytes:timeout:, TLS wrapClient:host:timeout:, and the TLS read / write helpers all integrate with the actor scheduler. A slow upstream parks the calling fiber on wait_readable / wait_writable instead of blocking the OS thread, so other actors keep running. Hostname resolution runs on a background resolver thread and the connect itself is non-blocking, so neither a slow DNS server nor a slow handshake freezes the domain either.

One rule holds across every socket, UDPSocket >> receiveFrom: included: a bounded read answers Maybe none at clean EOF and raises when the deadline the caller set passes. A timeout is caught with Error try:onError:; it never slips past an ifPresent:ifAbsent: as an absent value.

TCP listen: binds every interface; TCP listenHost: addr port: port binds one local address, as do HttpServer startHost:port:handler: and startGracefulHost:port:handler: one level up. Two servers can therefore hold the same port on different addresses, so a local run of several nodes keeps the port its deployment uses and separates the nodes by address instead.

Every connected socket sets TCP_NODELAY — accept, connect, the Remote client, the HTTP client, the Redis plugin — and a response is built in a buffer and written once, rather than as a status line, then each header, then the blank line, then the body. Nagle's algorithm holds a small write until the peer acknowledges the previous one, and the peer's stack delays that acknowledgement by around 40 ms, so any exchange that left in two writes stalled — which is the shape of every request/response protocol here: HTTP, JSON-RPC, TERIOS, grain. On Linux the fix took a raw TCP round trip from 88 ms to 18 µs and a Remote RPC call from 44 ms to 70 µs. Windows never hit the stall, which is why this survived until the deployment platform was measured.

The TLS bridge drives Tls.Engine directly — handshake, application read, and write each park on the runtime's I/O wait primitives.

Not on the Roadmap

Erlang / Akka Cluster style distributed actors — cross-node spawn, link, monitor, supervision over the wire — are not on the roadmap. The local actor model stays as-is.

Multi-node approach — virtual actors (grains)

The multi-node path ships as a virtual-actor (grain) model. A grain is an actor addressed by a stable (class, id) identity rather than a fiber reference, with its live state in an external store instead of pinned process memory. Because it is reached by identity, an activation can be dropped when idle and rehydrated on demand — there is no cross-node reference graph to garbage-collect. The application writes ordinary methods on the grain class and sends messages to the identity; the distributed lifecycle lives in the Grain module (stdlib/grain.lp).

The runtime upholds one-live-object-per-identity with three roles:

Single-activation holds unconditionally within a process — the directory's get-or-create cannot race. Across processes it holds under two preconditions: a single, linearizable Redis, and owners that do not stall past the lease TTL. The fence orders writes; it does not serialize two owners' read-modify-write. Outside those preconditions — a Redis failover that drops the newest token, a partition that leaves the old owner alive, a GC pause longer than the lease — two owners can briefly coexist and a concurrent update can be lost. Recovering from that is the application's job: write back to the system of record idempotently.

Single-node mode (for:initial:redis:) uses the directory and owner only; cross-node mode (for:initial:redis:url:) adds the lease. Routing is redirect-based and hidden by the client: a node that does not own an id answers with the owner's URL, and the Remote proxy follows the redirect transparently, caching the owner per id — application code keeps a single entry URL. With a static member list (for:initial:redis:url:members:) each id hashes to a preferred node, so activations spread over the cluster even when every first contact hits one entry node; if the preferred node is unreachable, the proxy retries the entry node with a takeover flag, and the entry node takes the id over once the dead node's lease expires. Redis is the volatile activation store, not a durable record — persisting to a system-of-record database is the application's responsibility.

Storage is pluggable. GrainServer talks to a GrainStore protocol (stateFor:, persist:json:token:, acquireFor:, renew:token:, release:token:). RedisGrainStore is the built-in backend; an application can inject its own (for example an actor-held in-memory map) with for:initial:store:.

Selective receive / link / monitor remain local-only by design — there is no cross-node link cascade. The remote ref Remote at: url for: #Class id: is typed when given a literal class: an unknown selector is a compile-time error (a non-literal for: stays dynamically typed as before).

A Remote call reuses its connection. Sockets are kept per host and port rather than dialled per call (up to eight per endpoint), and one the peer closed in the meantime is retried once on a fresh connection. On Windows, where no Nagle stall masked it, sequential JSON-RPC calls to one endpoint went from 4543 µs each (220 calls/sec) to 1411 ± 256 µs (709 calls/sec). Kept connections are not evicted by idle time, so a caller that reaches many endpoints sporadically holds their file descriptors.