Tree
- Tree:
14228021555c160076c41dbd76b74119b9049ee8- Date:
- Message:
- re-architecture: parser process, account worker, admission control Twenty-five changes, made and checked on the test machine one at a time, committed together. Each follows in the order it was made, under its own subject line, with its own account of how it was checked. Together: COPY and MOVE link messages instead of reading them into memory; SEARCH parsing moves out of its own process and message parsing into a parser-worker; each logged-in account gets one store child shared by its sessions, with a parser only while one is needed; the auth-worker exits after its grant; three admission limits bound what one account, one address and the whole daemon may hold; TCP keepalive ends the sessions of clients that vanished; and SEARCH answers SUBJECT, HEADER, SENTBEFORE, SENTON, SENTSINCE, FROM, TO, CC and BCC in the parser-worker. Change 1 of 25: COPY and MOVE link messages instead of reading them into memory COPY and cross-mailbox MOVE used to read every message in the range whole into the store child's memory before writing anything, bounded at 64 MiB per message and 512 MiB per command. With 64 store children that is a large amount of memory reachable by ordinary clients, and the per-message bound refused to copy messages the server had accepted: "append max" goes up to 1 GiB, and mail from an MTA is not bound by it. A maildir message is never rewritten once delivered; its flags live in its name. So the copy is now the same file under a new name: pass 1 records each source path and its new name and reads nothing, and pass 2 linkat(2)s each one straight into the destination's cur/. If the link fails for any reason other than EEXIST (another filesystem, the link count at LINK_MAX, a file flag), the message is streamed through a 64 KiB buffer into tmp/, fsynced and renamed, as before but without holding it whole. EEXIST fails the COPY instead of falling back, because the fallback's rename(2) would replace the file already there. AT_SYMLINK_FOLLOW keeps the old behaviour for a symlinked message, whose target was what openat(2) used to read. pledge(2) and unveil(2) already allow this: SYS_linkat is PLEDGE_CPATH, which the store holds, and its "rwc" unveil gives the read and create permissions dolinkat() checks. The existing rollback is unchanged and covers both paths. Cross-mailbox MOVE still commits the destination before unlinking the source, so it now moves no data. COPY_STAGE_MSG_MAX and COPY_STAGE_TOTAL_MAX are gone, as is the banner comment that said commit_copy_messages() had no rollback. It has one, and a reviewer had believed the comment over the code. stage_copy_messages() lost a directory descriptor parameter that disagreed with the global it also used. A new wire test APPENDs a message one octet over the old cap and checks COPY to another mailbox, COPY into the selected one, and MOVE, each by size and by octets at the start, middle and end. Against the old code it failed all three for the right reason, the cap, as maillog showed; after, 24 of 24. It needs "append max" and "attachment max" raised, so it is run by hand. The copy test's optional low-cap section is removed with the cap. Not exercised: the fallback copy. Nothing a client can do makes a link fail. Change 2 of 25: remove the per-connection SEARCH-parsing process Every connection forked a third process whose only job was to parse the SEARCH grammar and send the parsed program back over imsg. The parse now happens in the listener-worker, which already parses every other command in that same process, holds no private key, no password hashes and no mail, runs as _imapd in /var/empty and is pledged "stdio recvfd". A grammar bug there reaches one session, because every connection has its own listener-worker. The isolation that remains is the one that faces a stranger's data: message content is parsed by the store child, and moving that parse into a process without write access to mail is separate work. A connection before login now costs two processes rather than three. On a 2013 Celeron J1900 with 8 GB, each process costs about 1.07 MiB of machine memory whatever it does, measured by free page count across 0, 10, 20 and 40 connections. Gone with it: PROC_SEARCH, IMSG_SETUP_SEARCH_PEER, IMSG_SEARCH_PARSE_REQUEST and IMSG_SEARCH_PARSE_RESULT, SESSION_SEARCH_PARSING, iev_search, listener_dispatch_search(), the parent's per-connection fork, its reaping and its search_pid bookkeeping, and the _imapsearch account the role dropped to. An installation that created that account can remove it. search_dispatch() now parses and calls search_dispatch_finish() directly; both, and the parser, are static to search_cmd.c. The parser is search_program_parse(), named for what it does rather than for the process that used to call it. struct imsg_search_parse_result is struct search_parse_result, and the two constants lose their ORACLE infix, since neither crosses an imsg any more. Clients see no change: the same BAD and NO replies from the same parser, and the same 8192-octet bound on criteria, which is smaller than the command line it is sliced from. The two "[UNAVAILABLE] search temporarily unavailable" replies are gone; both existed only for a missing or dead oracle channel. The build and the test suite are clean on OpenBSD 8.0-current, and the lock-timeout test of SEARCH passes 5 of 5 with "lock timeout 5". Change 3 of 25: accept a MIME body part that declares no headers A multipart body part may carry no header fields at all. RFC 2046 section 5.1.1 says the boundary delimiter is terminated "by either another CRLF and the header fields for the next part, or by two CRLFs, in which case there are no header fields for the next part", and that a part with no Content-Type field is text/plain. find_header_body_split() looked only for two consecutive line breaks, which such a part does not contain: it begins with the second CRLF and then the body. The function returned -1 and both callers read that as malformed. build_body_structure() failed the whole structure, so FETCH BODYSTRUCTURE answered "* n FETCH ()" with a tagged NO; find_mime_part() found no part, so FETCH BODY[n] answered an empty string with a tagged OK. Such a message arrives through APPEND, so producing one needs no access to the server. The function now treats a buffer that begins with a line break as having an empty header section, before the scan for a blank line. A part that does declare headers is unaffected: the scan it relies on is unchanged and still runs. Checked on the test machine against a twelve-message corpus covering nested multiparts, a message/rfc822 part, a base64 attachment, body text that nearly matches the boundary, an empty body, a folded header and quoted Content-Type parameters. Of every FETCH result that corpus records, exactly two moved. The part with no headers now reports ("TEXT" "PLAIN" ("CHARSET" "US-ASCII") NIL NIL "7BIT" 26 0) inside its MIXED parent, and BODY[1] returns its 26 octets. Every other line was byte-identical, including the two BODYSTRUCTURE refusals that are deliberate: message/rfc822 is still scoped out, and a part specifier naming a part that is itself multipart still returns nothing. Change 4 of 25: add a parser-worker process beside each store child The store child parses hostile message content in the same process that holds the account's maildir. A bug in the MIME or RFC 5322 parsing of a message reaches the mail it is parsing. This adds the process that separates the two, and nothing else: it answers no requests yet, so the tree's behaviour is unchanged. parser.c is the new role. It drains its init message and its peer descriptor on fd 3, refuses to run as uid or gid 0, chroots to /var/empty, drops to the account's own uid, and pledges "stdio recvfd". No "rpath", so openat(2) is denied (kern_pledge.c) and the confinement is a property of the process rather than of the file's care. It opens nothing: every byte it will parse arrives on a descriptor its peer passes with the request. The parent forks one beside each store child, because the store child cannot fork its own: its pledge has no "proc exec". The parser is told the account's uid and given that store child's channel, and nothing else; it is told no paths because it opens no files. Its lifetime needs no bookkeeping. The store channel is its only one, so it reads EOF and exits whenever that child does, however the child dies. The one window between the fork and the pairing, where it would have no store to outlive, is closed by hand. A store child now has two peers, so setup_recv_one_peer() grew a variant that reports the imsg id. The listener's IMSG_SETUP_PEER carries the session number and the parser's carries 0, which is safe as a discriminator because next_session_id starts at 1 and wraps to 1, so no session ever carries it. The store sorts its peers by that id rather than by arrival order, and refuses a duplicate or a missing one. Checked on the test machine. The build and the test suite are clean, and the twelve-message FETCH corpus is byte-identical, as it must be when no request type exists. Over six live sessions, ps shows one parser per store child, each at its own store's uid and none at root. Their state flags read Ipc against the store's IpUc: pledged and chrooted, with no unveil at all, which is the intended shape showing up in the kernel's own accounting. The store carries a locked unveil because it opens the maildir; the parser has none to lock because it opens nothing. Change 5 of 25: build ENVELOPE in the parser-worker, and check what comes back This moves the first attribute behind the parser-worker added in the previous change. build_envelope() now runs there, on a descriptor the store opens and passes, so the RFC 5322 header parsing and address-list decomposition of a hostile message no longer runs in the process that holds the maildir. read_message_header() keeps its signature and becomes an open, a call to a new descriptor-taking core, and a close. The core is what the parser can run: it has no "rpath" and cannot open a message itself. build_envelope() takes the message descriptor in place of the mailbox directory descriptor, which is the same argument shape, so no declaration changed. The request and its reply share one imsg type, as the keymgr forwarders in listener.c do, with a per-request serial on the imsg id field and the reply checked on both type and id so a stale one is discarded rather than answered. The store's side of the channel is a bare imsgbuf rather than an imsgev: every exchange is synchronous, so a libevent dispatcher would only race the reply the child is already blocked waiting for. The wait is bounded, which the comparable wait on keymgr is not. The peer here is the one process in this design assumed to be attackable, so a parser looping on crafted MIME would otherwise pin its store child for the session's life, and store children are a fixed, machine-wide resource. relayd bounds the same shape of wait at one second (RELAY_TLS_PRIV_TIMEOUT, usr.sbin/relayd/relayd.h); this is longer because a parse is not one private-key operation. sshd's privsep client blocks on its monitor with no bound at all (usr.bin/ssh/monitor_wrap.c), but there the process waited on is the privileged one. A parser that misses the deadline once is not asked again, since retrying would let a stuck one charge the full timeout per message. What comes back is checked before it is used. ENVELOPE text is spliced into an untagged FETCH response as raw bytes, which was safe while the store built it and was trusting its own builder. It is not safe when another process supplies it: a CR or LF inside the reply ends the untagged line early and what follows is read as a new server response. RFC 9051 section 4.3 excludes NUL, CR and LF from a quoted string, and parser_reply_safe() refuses a reply carrying any of them, or one that does not close its own parenthesised list and would swallow what follows it on the line. A confined process whose output is trusted is not confined. A refused or missing reply is reported the way an unreadable message already is, so no wire shape and no listener code changed. A fuzzing harness drives the check, built two ways on purpose: without the check, as the tree stood before it existed, it must report violations. It found 243988 of them over 308466 cases, and none with the check in place, with 8531 replies still accepted as safe so the check is not passing by refusing everything. Writing its cases found a fault in the check before any of it shipped: the first draft walked the text once and skipped a backslash-escaped byte without examining it, so a quoted backslash followed by CR would have passed. The scan for line breaks is now a separate unconditional pass over every byte. Checked on the test machine. The build and the test suite are clean and the twelve-message FETCH corpus is byte-identical, which is the result that matters here: every one of those envelopes is now built in another process, on a descriptor rather than a directory, marshalled back and validated, and not a byte moved. Change 6 of 25: strip NUL, CR and LF from an address display name envbuf_append_nstring() has always substituted a space for NUL, CR and LF, citing RFC 9051 section 4.3, which excludes them from a quoted string, and substituting rather than rejecting so one bad byte does not drop the field. envbuf_append_one_address() open-codes a second quoted string emitter for the display name, because it also has to unescape the source's backslash escapes, and that copy was missing the substitution. Two emitters in one file, one of which sanitised. The mailbox and host parts go through the sanitising one and were never affected; only the display name was. So a From header reading From: "quoted<CR>name" <a@b> produced an ENVELOPE carrying that CR inside a quoted string, which went into an untagged FETCH response as raw bytes. Such a message needs no access to the server: APPEND literal octets reach the store unfiltered, so a client can store one itself. Only a bare CR reaches this far. A full CRLF does not: hdr_next_field() ends the field there, and a folding continuation does not carry it through. So the effect is a corrupted response line rather than a forged one, and since the previous change it is neither, because parser_reply_safe() refuses the text. That is the wrong outcome too: the message loses its ENVELOPE rather than its CR. The substitution now runs after the unescape, so a quoted backslash followed by CR is caught as well. vis(3) is how base makes untrusted text safe to emit, and mda.c wraps it in quotes exactly as a quoted string is wrapped (usr.sbin/smtpd/mda.c). It does not fit here. VIS_SAFE deliberately treats CR as visible and would have left this alone (lib/libc/gen/vis.c), and without VIS_SAFE every byte above 0x7f is escaped, which would mangle the UTF-8 that RFC 9051 admits in QUOTED-CHAR. The exclusion set is not a judgement call: TEXT-CHAR is any character except CR and LF, so those two and NUL are the whole of it and every other octet is preserved. A fuzzing harness over the ENVELOPE builder found this. Its property is not crash-freedom, which the sanitisers already cover, but that every envelope the builder can produce is one parser_reply_safe() will send. A gap between those two sets is not an attack, it is a message quietly losing its ENVELOPE. It reported three violations over 402003 cases, all this one cause, and none after the fix. This is the path the parser-worker exists to contain, and it had no harness at all before this one, which is where the headerless body part bug lived until a characterisation run found it. Checked on the test machine: the build and the test suite are clean and the twelve-message FETCH corpus is byte-identical, which is what it should be, since none of those messages carries a bare CR and the substitution must not fire on ordinary mail. Change 7 of 25: build BODYSTRUCTURE in the parser-worker too This moves the attribute the split was built for. build_body_structure() is the recursive RFC 2045/2046 walk bounded by MIME_MAX_DEPTH and MIME_MAX_PARTS, and it is where hostile MIME does its real work. It now runs in the parser, on a descriptor the store opens and passes. What MIME parsing the store still does, for BODY[<section-part>], moves in the next change, with HEADER.FIELDS. read_message_body() splits the way read_message_header() did in the previous change: it keeps its signature and becomes an open, a call to a new descriptor-taking core, and a close. BODY.PEEK[<section-part>] still uses the whole-message read from the store and is unaffected; it moves in its own change. build_bodystructure() takes the message descriptor in place of the mailbox directory descriptor. The parser reads whole messages now, so it needs the same ceiling the store has: bodystructure_read_max rides on IMSG_PARSER_INIT, as it already rides on IMSG_STORE_INIT. The buffer is still sized to the message rather than to the ceiling, so the cap costs nothing on ordinary mail, and the parser frees it before answering, holding no state between requests. The channel is now written once rather than per attribute. The request and reply structs lose their attribute-specific names, and parser_envelope() becomes parser_request(type, label, dfd, basename, maxlen, ...) called directly from each of the two sites, with no wrapper in between. That is sshd's shape: mm_request_send() and mm_request_receive_expect() take the message type as an argument and sixteen typed callers use them directly, with no intermediate layer (usr.bin/ssh/monitor_wrap.c). The label names the attribute for the log, as read_message_body()'s already does. The alternative was a second copy of the serial, the poll(2) deadline, the id check and the reply validation; this tree already has one duplicated request loop, keymgr_forward_rsa() and keymgr_forward_ecdsa() in listener.c, and the timeout missing from it is missing from both copies. BODYSTRUCTURE needs no new check on the way back. It is spliced into the untagged FETCH response as raw bytes two lines below ENVELOPE, so parser_reply_safe() already covered it. The ENVELOPE harness grew to drive the recursive builder before this change was made, so that a fault would surface on the old code rather than after the move: build_body_structure, parse_content_type, split_multipart, mime_read_token_or_qstring, mime_str_upper and mime_is_tspecial. It found nothing. Unlike the address path, every quoted field here goes through envbuf_append_nstring(), which has always substituted a space for NUL, CR and LF, so the fault fixed in the previous change never existed here. That the result is not vacuous was checked by hand: a bare CR in a Content-ID, a Content-Description or a Content-Transfer-Encoding does reach the builder and does come back substituted. Checked on the test machine. The build and the test suite are clean and the twelve-message FETCH corpus is byte-identical, which is the whole claim: build_envelope() and build_bodystructure() now have exactly one caller each, both in the parser, so every envelope and every body structure in that corpus was produced in another process, read from a descriptor, marshalled back and validated, and not a byte moved. The two deliberate refusals are intact: message/rfc822 is still scoped out, and a part specifier naming a multipart part still returns nothing. Over seven live sessions ps shows one parser per store child at its own store's uid, none at root, each pledged and chrooted with no unveil, and their resident sizes unchanged from before they read whole messages. Change 8 of 25: move BODY[<part>] and HEADER.FIELDS into the parser-worker These were the last two FETCH items the store still parsed. A section-part ran the whole MIME walk, parse_content_type() and split_multipart() included, and HEADER.FIELDS split RFC 5322 fields and their folds, both in the process that holds the maildir. Both now run in the parser on a descriptor the store passes. The store still looks for the blank line that ends a header, for BODY[HEADER], BODY[TEXT] and BODY[], as smtpd's queue does when it bounces a message with headers only (usr.sbin/smtpd/bounce.c), and does nothing else with message content. Neither item is spliced into the response line the way ENVELOPE is. Both go out as literals, whose length is framed, so a CR or LF from a hostile parser is only an octet. What such a parser could still do is put a NUL on the wire, which RFC 9051 section 9 allows only in literal8, or name a range the listener cannot read, which it answers by shutting the connection because the literal's length has already gone out. So each reply has its own check in the store, chosen by type in parser_request(): parser_literal_safe() header octets: no NUL, within FETCH_HEADER_MAX parser_extent_safe() a part's offset and length inside the file, with no overflow extent_has_nul() no NUL in that range, read by the store The part's octets never cross the channel. The parser returns where the part lies, and the store checks that against fstat(2) on its own descriptor, never against a size the parser reported. It is the same descriptor the listener then reads: the caller opens the message once and parser_request() sends a dup(2) of it, so the extent is checked against the file the client is actually sent. ENVELOPE and BODYSTRUCTURE open at the call site too, so all four callers use one shape. read_message_header_fields() splits into filter_header_fields(), which works on a header already in memory, and a descriptor-taking caller. extract_mime_part() takes the message descriptor in place of the mailbox directory descriptor. read_message_body() lost its last caller and is removed. A section-part that names a multipart or message/rfc822 part still comes back empty, exactly as a missing part does. That is a deliberate divergence, and this change keeps it. The harnesses and tests landed before the change: - The parser baseline fetches HEADER.FIELDS, HEADER.FIELDS.NOT, a partial part and a missing part. Before this, its reference had never seen HEADER.FIELDS at all. - The ENVELOPE harness drives find_mime_part() and locate_mime_part(). It pins which parts are found, including the divergence above, and checks that every part found lies inside its body. Both directed-case checks were mutation-tested: flipping one expectation makes the harness fail. After the split it also drives filter_header_fields(). A selection and its .NOT must together make up the whole header plus one blank line, the filter must be idempotent, and any header the store would have read must pass parser_literal_safe(). - The reply-check harness has directed and swept cases for the three new checks. Built without them it reports NULs in literals, an oversized literal and short reads; built with them it reports none, with 20288 header literals and 473 part extents still accepted. - The header-walk test's generator had gone stale: its stub no longer matched read_message_header_fields()'s signature, so it could not build against the tree. It is repaired and now splices filter_header_fields(). Its 24 rows are byte-identical to the committed test's, both before and after the split. Checked on the test machine. The build is clean with no compiler warnings, and the test suite passes, 16 scripts of 16. Before this change was installed, the extended FETCH corpus was recorded against the old build: its earlier items were byte-identical to the previous reference, and all 48 new ones answered OK. After installing, the corpus is byte-identical to that record, all 144 fetches. read_message_header_fields() and extract_mime_part() now have exactly one caller each, both in the parser, so every header selection and every part in that corpus was found in another process and checked by the store, and not a byte moved. The two deliberate refusals are intact: message/rfc822 BODYSTRUCTURE is still scoped out, and a part specifier naming a multipart part still comes back empty. ps shows the parser beside its store child at the store's uid, pledged and chrooted with no unveil, at the same resident size as before, and maillog has no refused reply and no missed deadline. Change 9 of 25: gather a session's FETCH walk, IDLE and APPEND state into one structure The store keeps everything about its one session in file scope. That is right while a store child serves exactly one session, and it is what stands between this tree and a process that serves several sessions of one account. This is the first step of undoing it, and it changes no behaviour. struct store_session (store_internal.h) now holds the listener's channel and four pieces of per-session state that were statics until now: the paced FETCH walk and its batch counters mbox_fetch.c the IDLE change probe index.c the IDLE baseline, the UIDs last reported index.c the one APPEND in flight mbox_manage.c store.c keeps the one instance and registers it as the listener channel's callback argument, so store_dispatch() receives it. It reaches the handlers that use this state as an explicit argument: handle_mbox_select(), handle_mbox_fetch(), handle_mbox_idle_refresh() and the three APPEND handlers, together with the fetch_walk_*(), idle_*_reset() and append_abort() helpers. The lock-wait machinery records the session rather than the channel, so a deferred command is replayed against the same state it was deferred with. No handler reaches the structure by name, so nothing in it can be read on behalf of the wrong session once a process holds more than one. Still in file scope, and moved by the next change: the selected mailbox and its descriptor, the cur/ snapshot, and the lock-wait record itself. session_id stays a process global; it labels log lines. A wire test was added first. It fetches forty bodies in one command, which makes the walk pause twice at FETCH_FD_MAX and resume from the store's EV_WRITE handler. Nothing in the suite had fetched more than sixteen bodies in one command, so the pause and the resume had never run under it. It is now in the suite. Checked on the test machine. The new test passed against the build before this change, so it pins existing behaviour, and passes against this one. The build is clean with no compiler warnings, and the test suite passes, 17 scripts of 17. The FETCH corpus is byte-identical to the reference, all 144 fetches. With a five-second lock timeout, a FETCH, SELECT and APPEND held behind another session's STORE each wait, are refused NO [INUSE] within the bound, and leave the session usable, and an IDLE whose seed waited reports the holder's change once the lock is free. With seven sessions over three accounts logged in, ps shows one parser beside each store child at the store's uid, and maillog has no refused reply and no missed deadline. Change 10 of 25: drop the cur/ listing of a command that finished after a lock wait A command reads cur/ once and keeps the listing for as long as it runs (cur_snap, mime.c), and store_dispatch() drops it after every request. A command that could not take the index lock is not finished there: it is re-run from a timer by deferred_retry(), and nothing dropped the listing that run built. A STORE renames the files whose flags it changes, so after a STORE that had waited, the listing named files that were gone. The session's next command then found the old name, failed to stat it, and skipped the message as "indexed but missing on disk". A FETCH left the message out of its answer. A second STORE on it answered OK and changed nothing. A name absent from the listing falls through to a real scan of cur/; a stale name present in it did not. deferred_retry() now drops the listing when the command it re-ran is done, under the same rule as store_dispatch(): not while a FETCH is paused mid-walk, since that command has not finished. A wire test was added first. It makes a second session's STORE wait behind a STORE 1:* over a large mailbox, then checks that the next FETCH answers for the message with its new flags, and that a following STORE takes effect. Like the lock-timeout test it needs a large mailbox, so it is run by hand. Checked on the test machine. Against the build before this change the new test failed 4 of its 12 checks: in both rounds the FETCH after the waited command gave no response for the message, and maillog logged it as missing on disk. In the second round maillog shows only the FETCH skipping it, not the STORE between (a STORE plans in two passes and would log twice), which fits that STORE having waited too and left a stale listing of its own. With this change the build is clean with no compiler warnings, the new test passes 12 of 12, and the suite passes, 17 scripts of 17. Change 11 of 25: move the selection, the cur/ listing and the lock wait into the session The second step of gathering a store child's per-session state into struct store_session, after the FETCH walk, IDLE and APPEND. It changes no behaviour. What moved, from file scope into the session: mailbox_selected, the store's own gate store.c selected_mailbox, the selection's name store.c mailbox_dir_fd, the selection's directory store.c cur_snap, one command's listing of cur/ mime.c deferred, the command waiting for a lock store.c The handlers that reach any of it now take the session in place of the listener's channel: STORE, EXPUNGE, SEARCH, STATUS, COPY, MOVE, DELETE and RENAME, with the helpers under them. SELECT, FETCH and the IDLE refresh already did. CREATE, LIST, SUBSCRIBE and UNSUBSCRIBE touch none of it and keep the channel. The cur/ listing functions in mime.c take the listing they work on, struct cur_snapshot, rather than a static, so locate_message_file(), open_message_file(), read_message_header() and message_body_range() gain it as their first argument. The parser links mime.c and calls none of them. The lock-wait record keeps its one slot, now per session. Its retry timer carries the session as the callback argument, as the listener channel already does, so deferred_retry() and the functions around it take the session instead of reaching a static. What is left in file scope is per process: session_id, which labels log lines, the maildir root, the configured limits, append_counter and the parser channel. The delete-selected wire test was extended first. After a RENAME of the selected mailbox, a COPY and a MOVE to the new name must take the same-mailbox path. They choose it by comparing the destination with the selection's name, which is one of the things that moved. Checked on the test machine. The build is clean with no compiler warnings, and the test suite passes, 17 scripts of 17, the extended delete-selected test 18 of 18. The FETCH corpus is byte-identical to the reference, all 144 fetches. With a five-second lock timeout, each of the eleven commands that can wait for the lock (STORE, EXPUNGE, FETCH, SEARCH, CLOSE, SELECT, STATUS, COPY, MOVE, APPEND and IDLE) waits behind another session's STORE and is bounded, refused NO [INUSE] and followed by a working session, or for IDLE reports the holder's change. The lock-wait snapshot test passes 12 of 12. With seven sessions over three accounts logged in, ps shows one parser beside each store child at the store's uid. maillog has no refused reply and no missed deadline, and nothing has been logged as missing on disk since the lock-wait fix went in. Change 12 of 25: attach a store child's session at runtime; the parent says when it exits A store child used to receive its one listener channel during setup, serve it, and exit when the session ended. It now takes its sessions at runtime over its parent channel, as keymgr takes listener peers, and exits only when the parent tells it to. The parent still attaches exactly one session to each child, so nothing a client can see changes; every login now goes through the attach path, and every logout through the detach path, before any child carries a second session. In the store child (store.c): setup brings only the parser peer, then IMSG_SETUP_DONE store_parent_dispatch() keeps fd 3 on an event and takes IMSG_SETUP_PEER (attach a session, its number in the id) and IMSG_STORE_EXIT (detach everything and exit) store_attach() allocates the session and opens its directory; a duplicate id, or id 0, is refused store_detach(), on IMSG_STORE_SHUTDOWN or the listener's channel closing, frees one session and leaves the child running the sessions are a list, not a static pledge gains recvfd, the set smtpd's queue pledges The handshake can read the first attach along with IMSG_SETUP_DONE, so the dispatcher is run once by hand before the event loop starts; without that the session would sit in the buffer until the parent next wrote. If the parent goes away, the child serves the sessions it has and exits with the last of them, which is what it did before. In the parent (parent.c), the listener's channel is sent after IMSG_SETUP_DONE rather than before the parser's, and reap_child() sends IMSG_STORE_EXIT to a session's store child when it reaps that session's listener. IMSG_STORE_EXIT is new in imapd.h. A wire test was added first. It opens and ends 80 sessions one after another, half with LOGOUT and half by dropping the connection, and checks that every one logs in. A store child that outlived its session would count against STORE_CHILD_MAX (64) until logins were refused. It is in the suite, now eighteen scripts. Also, a comment in store_dispatch() that pointed at mime.c for the cur/ listing now points at struct cur_snapshot, where that description moved. Checked on the test machine. The new test passed against the build before this change, and passes against this one. The build is clean with no compiler warnings, the suite passes, 18 scripts of 18, and the FETCH corpus is byte-identical to the reference, all 144 fetches. The lock-wait snapshot test passes 12 of 12. ps lists the same processes before and after the store-exit test's 81 sessions, and with a mail client logged in to three accounts it shows one parser beside each store child at the store's uid. maillog has no refused reply, no missed deadline and no refused session. Change 13 of 25: one store child per account: later logins join it The parent now keeps one store child per account, keyed by uid, gid and maildir, and attaches each later login of that account to it over the parent channel, the way the first one is attached. The child serves the account's sessions side by side, each with its own selection, FETCH walk, IDLE state and lock wait, as the last three changes arranged. Twelve connections of one account used to cost twelve store children and twelve parsers; they now cost one of each. In the parent (parent.c): struct store_child records the account it serves, how many of its sessions' listeners are alive, and whether it has been told to exit; open_session records its store child parent_handle_store_fork() attaches to a live child of the same account, or forks one if there is none store_child_session_gone() sends IMSG_STORE_EXIT when the last listener goes, and a child told to exit takes no new login store_child_teardown() fails every session attached to a child that dies during setup, not only the first STORE_CHILD_MAX now counts accounts In the store child (store.c), a read, write or framing failure on one session's channel detaches that session rather than calling fatal(), since the account's other sessions are served by the same process. session_id, which labels log lines, is set to the session being served on every entry from the event loop. The store child's title is now "store uid N" and the parser's "parser uid N", since neither serves one session any more. One behaviour does change. A long command in one session, such as a STORE over a large mailbox, now delays every other session of the same account until it finishes, because one process serves them all. Before, only sessions that needed the same mailbox's lock waited, and a wait was bounded by "lock timeout" and answered NO [INUSE]. Between sessions of one account the lock wait no longer happens at all; it still guards against a second process. The parser is still one per store child, so it is now shared by the account's sessions, and a parser that misses its deadline or dies stops message parsing for the whole account until its sessions end. Starting it on demand and bounding its CPU is the next change. A wire test was added first. Two sessions of one account each see the other's APPEND and STORE, including while idling; one pages a 40-message FETCH through its pause while the other runs a command; one EXAMINEs INBOX without moving the other; and each goes on working after the other logs out or drops its connection. It is in the suite, now nineteen scripts. Checked on the test machine. The new test passed 14 of 14 against the build before this change and passes 14 of 14 against this one. The build is clean with no compiler warnings, the suite passes, 19 scripts of 19, and the FETCH corpus is byte-identical to the reference, all 144 fetches. With a five-second lock timeout, the lock-timeout test now fails its "bounded" and "NO [INUSE]" checks, as intended: the waiter is served after the holder's 47-second STORE, with OK, because both sessions are in one process. It passed before. ps shows two sessions of one account behind a single store child and a single parser, and six Mail.app sessions over three accounts behind three of each, where there were six. maillog has no refused reply, no missed deadline and no refused session. Change 14 of 25: parser on demand, with a CPU bound and three strikes per message With one store child per account, the parser became one per account too, and a parser that missed its deadline or died stopped message parsing for every session of the account until the last one ended. The parser is now started when a request needs one, let go when it is idle or used up, and replaced when it fails. In the store child (store.c, mbox_fetch.c), there is no parser at setup. A FETCH walk that needs one and has none asks the parent (IMSG_PARSER_WANT) and waits at that message, as it already waits for its queue to drain. It goes on when IMSG_SETUP_PEER with id 0 brings a parser, or without one on IMSG_PARSER_NONE or after PARSER_REPLY_TIMEOUT_SEC. Before each message the walk checks the parser's channel without blocking, so a parser that has died since its last request is replaced rather than blamed. A deadline missed, or the channel lost once a request has been sent, lets the parser go and counts against that message. At PARSER_STRIKES (3) the message is not sent to a parser again while the child lives; the count is kept in the child, never in the parser. A parser unused for PARSER_IDLE_SEC (60) is let go by closing its channel. The parser-channel failures that called fatal() now let the parser go instead, since the far end of that channel is the process assumed to be attackable. In the parent (parent.c), IMSG_PARSER_WANT forks a parser for that store child, ending any earlier one it still has, so there is at most one per account. RLIMIT_CPU is set between fork and exec, 8 seconds soft and 10 hard, because the parser cannot set it under pledge. In the parser (parser.c), a reply sent once it has used PARSER_CPU_RETIRE_SEC (4 s, half the soft limit) of CPU says so, and the parser exits after sending it; the store lets it go on reading that. Honest work therefore never reaches the limit. The worst honest request measured 1.7 s of CPU on the test machine. A wire test was added first. It fetches ENVELOPE and BODYSTRUCTURE, waits while the account's parser is killed on the server, and fetches them again. It needs a person on the server, so it is run by hand. Checked on the test machine. The new test fails 3 of 8 against the build before this change: once the parser is killed, FETCH answers neither ENVELOPE nor BODYSTRUCTURE, and says NO. It passes 8 of 8 against this one, with nothing new in maillog. The build is clean with no compiler warnings, the suite passes, 19 scripts of 19, and the FETCH corpus is byte-identical to the reference, all 144 fetches. ps, sampled every five seconds through one session, shows no parser after login, one from the first FETCH ENVELOPE, none again about a minute later with the session still open, and the store child gone at LOGOUT. maillog has no refused reply, no missed deadline, no refused session, no parser that failed to arrive and no message struck out. Change 15 of 25: hold a mailbox's index lock from outside imapd, for the lock tests One store process now serves every session of an account and runs one command at a time, so two sessions of one account never contend for an index lock: the second session's command is served after the first and never meets the lock. The two lock tests made their holder a second session of the same account, so they no longer reached the code that waits for a busy lock, which is kept for a second process. A small program added here is that second process. Run on the server as root and given an account's maildir and a mailbox in it, it takes flock(2) LOCK_EX on the mailbox's imapd.index.lock while a mailbox named lockhold exists at the top of the same maildir, and releases it when that goes. The tests CREATE and DELETE lockhold from a third session, so the test decides when the lock is held, and nothing on the two machines has to be started at the same moment. Taking one maildir for both means the lock and the signal cannot name different accounts. A hold ends after -t seconds (default 60) whatever the signal says, so a test that dies cannot leave the mailbox locked. It only reads: the lock file is opened O_RDONLY without O_CREAT, flock(2) takes LOCK_EX on a read-only descriptor, and it pledges "stdio rpath flock". The lock-timeout test can use either holder, the external one by default. The imap holder is kept: against the shared process it shows that two sessions of an account are served one after the other. With the external holder it makes 4 checks rather than 5, since the hold cannot be seen from the client; "not early" and "refused" show it together. IDLE is not tested with the external holder, which changes nothing: "+ idling" is sent before the seed, so a seed that met the lock would leave nothing on the wire to check. The lock-wait snapshot test uses only the external holder, since the imap one would now pass without reaching the lock-wait code. The hold is 2 seconds, so it runs under the same short "lock timeout" as the other test, on any mailbox holding the message it works on. It makes 10 checks rather than 12: the holder's own STORE is gone, and the check that the waiter really waited now times it against the release. Its docstring also named the cur/ listing by its place before it moved into struct cur_snapshot. No daemon code changes. Checked on the test machine, against the installed account-worker build with "lock timeout 5" and a mailbox of 100,000 messages. The holder builds with no compiler warnings. With it holding the mailbox's lock from outside imapd, the lock-timeout test passed 4 of 4 for each of the ten commands it sends (store, expunge, fetch, search, close, select, status, copy, move, append): each waited, was refused NO [INUSE] within the bound and not early, and left the session working. The holder logged ten holds of 6.3 to 6.5 seconds, none of which had to wait for imapd. That run used an earlier build of the holder that took the lock file and the signal as two paths; the rest used this one. The snapshot test passed 10 of 10. With the holder stopped, the timeout test failed 2 of 4, the two checks that need a held lock, so those checks depend on it. maillog has no "missing on disk", "cannot be sent", "did not answer within" or "refusing a session" line. Change 16 of 25: end the auth-worker once the parent holds its grant Each connection has its own auth-worker. Once it has granted a login it has nothing left to do: it refuses a second grant for its session, and IMAP has no way to authenticate again after login. It stayed alive only until the listener's channel closed at the end of the session, so every logged-in connection kept one process that did nothing. The parent now sends IMSG_AUTH_EXIT once it has accepted the worker's IMSG_AUTH_CRED, and the worker flushes its answer to the listener and exits. The parent says when, rather than the worker exiting as soon as it has answered, because reap_child() closes a child's channel without reading it: a worker that exited on its own could take with it a grant the parent had not yet read. The worker's end matches sshd's, whose pre-authentication child exits once authentication is done, and the parent deciding matches IMSG_STORE_EXIT. The listener now expects its auth channel to close after a login and logs that at debug level. Before a login it is still a warning. A connection that never logs in is unchanged: its auth-worker takes every attempt, and exits when the listener goes. One path changes as a consequence. After a login whose store spawn failed, a retried AUTHENTICATE reached a worker that refuses an already granted session without answering, so the client waited until the login grace ended the session, or for ever with "login grace 0". It is now answered NO [UNAVAILABLE], or, if the listener has not yet seen the worker go, the listener's write to it fails and the session ends. A wire test was added to the suite first. It checks that a failed attempt leaves the next one able to succeed on the same connection, that the session works once the worker is gone, that AUTHENTICATE and LOGIN after login are refused BAD, and that twenty logins in a row each succeed. A client cannot see the change, so it passes before and after; ps(1) on the server shows the difference. Checked on the test machine. Against the build before this change, the new test passed 8 of 8, and ps(1), sampled every 5 seconds, showed the held session's auth-worker alive for the whole two minutes it was logged in. With this change the build is clean with no compiler warnings, the suite passes, 20 scripts of 20, with the new test at 8 of 8, and the parser baseline is byte-identical to the reference. ps(1) over four minutes showed no auth-worker behind any logged-in session, the held one or the four others logged in at the time. maillog has no "auth-worker closed channel", "refusing IMSG_AUTH_CRED", "IMSG_AUTH_EXIT", "imsgbuf_flush IMSG_AUTH_RESULT", "cannot be sent", "did not answer within" or "refusing a session" line. Change 17 of 25: imapd.8: one store child per account, and what lock timeout bounds Three statements in the manual had gone stale. The description said a store child is forked per authenticated session. One store child now serves each logged-in account, shared by all of that account's sessions. "login grace" counted three processes per accepted connection, including the search-oracle, which is gone. A connection costs two until it logs in: its listener-worker and its auth-worker. "lock timeout" said a command waits for another session of the same user. That user's sessions share one store child, which runs one command at a time, so they never wait for each other's lock; a long command delays the account's other sessions until it is done, and the directive does not bound that. What it bounds is a wait on a lock held by another process: a store child whose sessions have all ended, still finishing a command, while a new login for the same account starts another. Checked on the test machine. mandoc -Tlint reports only the two STYLE notes it reported before: the missing RCS id, and imapduser(8) not found. The build it was checked on also carries the "account sessions" change. Change 18 of 25: limit the sessions one account may hold: "account sessions" Nothing bounded how many sessions one account could open. Each costs a listener process, all of them share the account's store child, and a client configured to open many connections could have as many as it asked for. The new directive "account sessions", default 8, caps them. An account is the store child's key, a uid, gid and maildir from the credentials file, and the parent already counts the sessions attached to each. When a login would pass the cap, the parent refuses to attach it, and the listener answers NO [LIMIT] (RFC 9051 section 7.1) and closes the connection. It does not keep the connection for a retry: the connection's auth-worker has spent its one grant, as sshd's pre-authentication child does, and sshd likewise disconnects when a limit is reached during authentication. No BYE is sent, since RFC 9051 section 7.1.5 lists four conditions for one and a refused login is not among them. The client may log in again on a new connection once one of the account's sessions has ended. The other failures on that path, a store child that could not be started or did not finish its setup, and a second grant for a session already authenticated, now close the connection the same way, after the same NO as before. Left open, such a connection could not log in again, and it was no longer counted by the startups throttle, which counts only connections not yet authenticated, so it lasted until the login grace ended it, or for ever with "login grace 0". The directive is reloaded on SIGHUP and documented in imapd.8 and imapd.conf.example. A wire test came first. It logs one account in up to the cap, and checks that one more login is refused NO [LIMIT] and closed without a BYE, that the sessions already open still work, and that once one logs out the next login succeeds and the one after it is refused again. It needs an account no other client is logged in as, so it is run by hand. Checked on the test machine. Against the build before this change the new test failed 4 of 7, as predicted: the ninth login and the one after the logout were both answered OK. With this change the build is clean, with no compiler warnings and no yacc conflicts, mandoc reports only its two existing STYLE notes, the suite passes, 20 scripts of 20, the parser baseline is byte-identical to the reference, and the new test passes 7 of 7. maillog shows each of its two refusals as the parent's "refusing login: uid ... already has 8 sessions (account sessions)" followed by the listener's "closed ... reason=limit-exceeded". Change 19 of 25: limit the connections held at once: "connections max" Nothing bounded how many connections imapd accepted. Each costs a listener process, and one more until it logs in, and the parent and keymgr keep a descriptor for each. On OpenBSD's defaults, the process table and the parent's descriptors run out at about a thousand connections, long before memory does, and would run out in the middle of spawning a session. The new directive "connections max", default 256, bounds them. The parent counts every session it tracks, logged in or not, and checks the count in its accept loop before MaxStartups and before any fork. Past the limit a connection on the cleartext port gets RFC 9051's rejected-connection greeting, "* BYE [UNAVAILABLE]", written best effort as sshd writes its own refusal, and is closed. On the TLS port that greeting would have to travel inside TLS, which the parent does not speak, so the connection is closed with nothing sent, as httpd does at its client limit. The parent logs once when the limit is reached and once when it accepts again, as smtpd does. C connections, L of them not yet logged in, for A accounts need up to C + L + 2A processes plus two, and the parent keeps a descriptor to each. The default leaves room for the worst case, every connection a different account, within OpenBSD's default kern.maxproc and the daemon login class's openfiles. imapd.8 gives that rule, and says to raise those limits before raising this one. STORE_CHILD_MAX goes. It capped store children, one per account, at 64, so it refused the 65th account to log in, well inside the new default, and "connections max" bounds the accounts anyway. accept(2) failing for descriptors now pauses the listening sockets for a second, as httpd does. Before, the listen event stayed armed and fired again at once. The comment on the parent's duplicate-grant path said it was reachable by an ordinary retry after a failed store spawn. It has not been since the auth-worker began exiting after its grant; only a broken auth-worker gets there. The directive is reloaded on SIGHUP and documented in imapd.8 and imapd.conf.example. A wire test came first. It needs a low "connections max" in the running imapd.conf, so it is run by hand. It opens connections until one is refused, and checks the refusal on port 143 is the BYE greeting and a close, the refusal on port 993 a close with nothing sent, that the connections held still work, and that once one logs out a new connection is accepted and the one after it refused again. Checked on the test machine. Against the build before this change the new test failed 4 of 6, as predicted: nothing was refused, and port 143 answered with its OK greeting. With this change the build is clean, with no compiler warnings and no yacc conflicts, mandoc reports only its two existing STYLE notes, the suite passes, 20 scripts of 20, and the parser baseline is byte-identical to the reference. With "connections max 6" the new test passes 6 of 6, itself holding 3 connections with 3 already open, and maillog shows "connections max reached (6 open), refusing new connections", then "below connections max again (5 open), accepting", then the first line again. Change 20 of 25: cap connections not yet logged in per address: "startups per-source" MaxStartups counts the connections not yet logged in across all addresses. One address could therefore hold every one of its slots, a hundred by default, with bare TCP connections and no password, and every other client was refused until the login grace closed them. The new directive "startups per-source", as sshd's PerSourceMaxStartups, caps what one address may hold, default 5, below "startups begin" so one address cannot even start the ramp; "none" turns it off. The parent keeps each session's address and counts, at accept and before any fork, the sessions from that address not yet logged in, as sshd's srclimit_check_allow() does. A session stops counting once it logs in. Each address counts on its own, IPv6 ones included, as sshd's default grouping does; imapd.8 says what that leaves open for IPv6, and that users behind one NAT address share the limit. sshd refuses PerSourceMaxStartups and MaxStartups through one function. All three admission limits now share one refusal here too: "connections max", the new cap, and MaxStartups, which until now closed silently on both ports. On the cleartext port a refused connection gets RFC 9051's rejected-connection greeting, "* BYE [UNAVAILABLE]", best effort, and is closed. On the TLS port it is closed with nothing sent, since the greeting would have to travel inside TLS. The directive is reloaded on SIGHUP and documented in imapd.8 and imapd.conf.example. A wire test came first, and joins the suite, which now runs twenty-one scripts, since the cap is on by default. From one address it holds the cap's worth of connections without logging in, and checks that one more is refused on each port in its own way, that the connections held still work, that one of them logging in makes room for another, that the next is then refused again, and that one logging out makes room too. Checked on the test machine. Against the build before this change the new test failed 3 of 7, as predicted: one more connection on each port was accepted. With this change the build is clean, with no compiler warnings and no yacc conflicts, mandoc reports only its two existing STYLE notes, the suite passes, 20 scripts of 20 before the new test joined it, the parser baseline is byte-identical to the reference, and the new test passes 7 of 7. maillog has no line for its refusals, which are logged at debug level, as MaxStartups' are. Once it had joined, the suite passed 21 scripts of 21. Change 21 of 25: set SO_KEEPALIVE on client connections After login nothing ended a session but its client. A client that vanished without closing its connection, a laptop put to sleep or a phone that lost its network, left its session open until the server next wrote to it, and a session that was never written to stayed for ever. With "account sessions" that is also a lockout: every such session counts against the account's limit, so a client that dropped off often enough could fill it until imapd was restarted. The parent now sets SO_KEEPALIVE on every connection it accepts, as sshd does by default for TCPKeepAlive, and as syslogd does for its TLS connections, so the kernel probes a quiet connection and drops it when the probes go unanswered. How long that takes is the kernel's: the net.inet.tcp.keepidle sysctl plus eight net.inet.tcp.keepintvl intervals, 7200 and 75 seconds by default, so 2 hours 10 minutes. imapd.8 says so under "account sessions". There is no directive: a connection is always probed, as syslogd's are. A failure to set the option is logged and the connection goes on. No inactivity autologout is added. RFC 9051 section 5.4 would allow one of at least 30 minutes, but once keepalive ends the sessions whose client has gone, all it would add is logging out clients that are alive and quiet, each of which then reconnects. A wire test cannot make its own client vanish, so the check is ktrace. Checked on the test machine. Against the build before this change, ktrace(1) on the parent while one connection was made showed accept(2) and the two fcntl(2) calls that set O_NONBLOCK, and no setsockopt(2). With this change the build is clean, mandoc reports only its two existing STYLE notes, the suite passes, 21 scripts of 21, the parser baseline is byte-identical to the reference, and the same ktrace shows "setsockopt(9,SOL_SOCKET,SO_KEEPALIVE,...,4)" returning 0 after the fcntl(2) calls. Change 22 of 25: parse the SEARCH content keys' operands before refusing them RFC 9051 section 6.4.4 lists eleven search keys that need a message's header or body: BCC, BODY, CC, FROM, HEADER, SENTBEFORE, SENTON, SENTSINCE, SUBJECT, TEXT and TO. imapd refused each by name with NO as soon as it read the key, so it never read the operand that follows. The listener now parses those operands, as the first step to answering the keys. The string keys take an RFC 9051 section 9 astring: an atom or a quoted string, with only the two quoted-specials escaped. HEADER takes a field name and a string. The three SENT keys take a date, read by the same function as BEFORE, ON and SINCE, now shared. Having parsed the whole program, SEARCH still answers NO for any of the eleven, with the same text as before, because nothing yet answers them. Two answers change now, because they can be given without matching: A literal operand is refused BAD, "literals are not supported in SEARCH, send a quoted string". A synchronizing literal is never sent its continuation, which section 2.2.1 allows for a command refused with BAD. Literals in commands other than APPEND are a change of their own. The string operands of one SEARCH may hold 4096 octets in all, counted after unquoting and including HEADER's field name. More is refused "NO [LIMIT]", the section 7.1 response code for an implementation limit. 4096 is section 4.3's cap on a non-synchronizing literal, and it leaves room for the operands beside the search program in one imsg. A malformed operand, such as an unterminated quoted string or a missing one, is now BAD rather than the old NO. Checked on the test machine. The build is clean and mandoc reports only its two existing STYLE notes; the suite passes, 21 scripts of 21; the parser baseline is byte-identical to the reference. A new wire test of the header keys, written before this change, makes 38 checks: before it 30 failed, each on the old NO; after it 27 fail, the same refusal, and the three that pass are the two over-cap SEARCHes answered NO [LIMIT] and the literal operand answered BAD with no continuation. Change 23 of 25: answer SEARCH SUBJECT, HEADER and SENT* from the message header RFC 9051 section 6.4.4's SUBJECT, HEADER, SENTBEFORE, SENTON and SENTSINCE are now answered, where before they were refused with NO. FROM, TO, CC, BCC, BODY and TEXT are still refused as before. The matching runs in the parser-worker, the process that already parses message content for FETCH, confined by pledge "stdio recvfd" in a chroot: the store never reads a header for SEARCH. A new request, IMSG_PARSER_SEARCH, carries the message's descriptor and the SEARCH's content keys with their strings; the reply carries one byte per key, 1 where the message matches it. The store checks that the reply is exactly that, one byte of 0 or 1 per key, before using it. The strings travel from the listener in a pool beside the search program, in IMSG_MBOX_SEARCH's fixed part, so the lock-wait record keeps them with the request as it is. A node names its string by offset and length, and the store refuses one that lies outside the pool. The store evaluates each message over three values first. A content key is unknown until the parser answers, and AND, OR and NOT carry that through, so a message its flags already decide is never sent to the parser: UNSEEN SUBJECT x asks only about unseen messages, and a SEARCH with no content key never starts a parser. The SEARCH walk can now pause, as the FETCH walk can: to wait for a parser to start, and every 64 parser requests to let the event loop serve the account's other sessions. A message the parser cannot check, because it missed its deadline, died, has struck out, or cannot read the header within "attachment max", ends the SEARCH NO. Answering as if it had not matched would look complete and not be. The matching itself, in search_match.c: A field's value is unfolded (RFC 5322 section 2.2.3) and its RFC 2047 encoded-words are decoded, B and Q, with the white space between two adjacent words dropped (section 6.2). The decoded octets are compared as they are: no charset is converted, since base has no iconv. Matching is a substring match, case folded in the ASCII range only, as RFC 9051 section 6.4.4 says it SHOULD be. HEADER matches any field of that name; an empty string tests that the field is present. SENT* take the Date: field's own calendar day, disregarding its time and zone, and accept RFC 5322 section 4.3's two- and three-digit years. A message with no Date:, or one that cannot be read, matches none of them. The header is read in blocks up to its blank line. A NUL in it does not stop the match, which goes by length. mime.c gains header_next_field(), which walks a header's fields for a caller outside the file, over the existing hdr_next_field(). parser_request() sends with imsg_composev(), as relayd's ca.c does, so a request can carry data after its fixed part; the four FETCH callers pass none. Checked on the test machine. The build is clean and mandoc reports only its two existing STYLE notes; the suite passes, 21 scripts of 21; the parser baseline is byte-identical to the reference. The header-key test, written before the grammar change, makes 38 checks; before this change 27 failed, after it 16, each a check that names FROM, TO, CC or BCC, still refused. SUBJECT "fill" matched all 10,000 messages of a bench mailbox in 2.4 s and 99,999 of 100,000 in 44 s; NOT SUBJECT "fill" matched the other one, whose Subject is a test's own. During the 100,000-message search, a second session of the same account logged in, selected and fetched in 0.36 s. ps showed a parser for the account during the search, replaced once as it retired, and maillog has no failed check, bad request or parser failure. Change 24 of 25: answer SEARCH FROM, TO, CC and BCC RFC 9051 section 6.4.4's FROM, TO, CC and BCC are now answered, where before they were refused with NO. Only BODY and TEXT are still refused. Section 6.4.4 matches these keys against "the envelope structure's" field, and an address in that structure keeps its display name, local part and domain apart (section 7.5.2). Each address is matched twice: as its display name, with its quoted-pairs undone and its RFC 2047 encoded-words decoded, and as mailbox@host rejoined, as smtpd rejoins an address as user@domain before matching a rule (ruleset.c, mailaddr_to_text()). So a whole address matches as it is written, and so does any part of it or of the name. Comments are not taken out: one after an address is matched as part of its host, and one in a display name as part of the name, as ENVELOPE shows them. What does not match: the angle brackets, the commas between addresses, and a group's name, since a group is no address. HEADER From still matches the field's text as it stands. The address parse is the one ENVELOPE already uses. It is split out of envbuf_append_one_address() as address_split(), and the top-level comma split out of envbuf_append_address_list() as address_list_next(); both formatters now call them, and what they produce is unchanged. The matching runs in the parser-worker, as SUBJECT's does, in search_match.c. Checked on the test machine. The build is clean and mandoc reports only its two existing STYLE notes; the suite passes, 21 scripts of 21; the parser baseline, which fetches ENVELOPE, is byte-identical to the reference. The header-key test makes 38 checks; before this change 16 failed, each naming FROM, TO, CC or BCC, and after it none. FROM "filler@example.com" matched all 10,000 messages of a bench mailbox in 2.4 s, as SUBJECT did, and 99,999 of 100,000 in 44 s. A FROM that matches nothing took the same 2.4 s, and SEEN FROM on a mailbox with nothing seen took 0.52 s, the time of a flag-only search, so no message went to the parser. maillog has no failed check, bad request or parser failure. Change 25 of 25: check the SEARCH matcher against references, and fix what that found search_match.c, the parser-worker's matcher for SUBJECT, HEADER, SENT*, FROM, TO, CC and BCC, now has a fuzzing harness. A sanitizer finds an overrun, not a wrong answer, so it also checks the answers against references written in the harness: ASCII-only case folding, RFC 2047 B and Q words decoding back to their octets, with the white space between two adjacent words dropped (section 6.2) and kept elsewhere, Date: days computed without timegm(), every obsolete year of RFC 5322 section 4.3, SENTBEFORE and SENTSINCE partitioning the dated messages, every mailbox@host of a generated address list, and headers read across the 64K block edge and the cap. A table carries the header-key test's messages and keys, and the address cases. Before the fixes below it failed, each time for a wrong answer: A word that fails to decode overwrote the space before it. decode_value() took the space back before decode_word(), which writes as it goes, so "=?x?Q?a?= =?x?B?QQ!A?=" decoded to "aA=?x?B?...". The word is now decoded after the space and moved back over it only if it decodes. A folded Date: was no date. RFC 5322 section 3.3 allows FWS between a date's tokens, and date_field_day() took only SP and HTAB. It now takes CR and LF too; the value's only CR and LF are folds. "attachment max" was rounded up to the next 64K block, and a message of header fields alone exactly that long was refused. read_header() now reads at most one octet past the cap and refuses a header longer than it, as read_body_from_fd() refuses a message longer than it. The header-key test joins the suite, now twenty-two scripts. Checked on the test machine. The build is clean; the suite passes, 22 scripts of 22, the header-key test among them; the parser baseline is byte-identical to the reference; the header test makes 38 checks and none fails. The matcher's harness, built there with the minimal UBSan runtime, ran 3,201,074 cases and exited 0; before the three fixes it failed on each of them, in a run elsewhere. maillog has no failed check, bad request or parser failure.
| README.md | commits | blame |
| contrib/ | |
| src/ | |
README.md
# OpenIMAPD
A from-scratch IMAP4rev2 ([RFC 9051](https://www.rfc-editor.org/rfc/rfc9051)) server for OpenBSD, written in C in the privilege-separated tradition of `smtpd(8)`, `httpd(8)`, and `ntpd(8)`. No third-party IMAP library.
**Status:** Pre-release, actively developed. Not a port. See [Getting the source](#getting-the-source) below for the repository.
## What it is
- **Privilege-separated**, `smtpd`-style, across six processes: a root *parent* reads configuration and binds the listening sockets; unprivileged *listener* and *auth* children handle the network and credential checks; *keymgr* holds the TLS private key and performs every private-key operation on request, so the process terminating TLS never has the key in its address space; a *parser* child, spun up by a *store* child on demand and retired after a spell idle, does every parse of attacker-reachable content — `SEARCH`'s grammar and `FETCH`'s `ENVELOPE`/`BODYSTRUCTURE`/header-field parsing alike — confined to `pledge(2)` `"stdio recvfd"`: it opens nothing itself, reading each message over a descriptor the store child passes it already open; and a *store* child is forked per authenticated account, shared by all of that account's sessions, chroots into the mail spool, and drops privileges to that account's own user before ever touching a message. `pledge(2)`, `unveil(2)`, and `chroot(2)` enforce these boundaries, not just convention: *parent* is the only process that can pass a file descriptor at all.
- **Storage**: stock maildir format (`tmp/`/`new/`/`cur/`, atomic delivery via `rename(2)`), readable with `ls` and `grep`, and natively understood by `smtpd(8)`'s own `maildir` delivery action. IMAP's extra bookkeeping (UIDs, UIDVALIDITY, per-message mod-sequences, keywords) lives in a small, `flock(2)`-guarded, line-oriented index file per mailbox, plain colon-delimited text, not a database.
- **Transport**: STARTTLS on port 143 and implicit TLS on port 993 ([RFC 8314](https://www.rfc-editor.org/rfc/rfc8314)), via `libtls`. `AUTH=PLAIN` only, refused before TLS is established.
## Protocol coverage
`CAPABILITY`, `STARTTLS`, `AUTHENTICATE`, `ID`, `ENABLE`, `SELECT`, `EXAMINE`, `CREATE`, `DELETE`, `RENAME`, `LIST`, `LSUB`, `NAMESPACE`, `STATUS`, `FETCH` (including `ENVELOPE`, `BODYSTRUCTURE`, and MIME-part-addressed `BODY[<part>]`/`BODY.PEEK[<part>]`), `STORE`, `SEARCH`, `APPEND`, `COPY`, `MOVE`, `EXPUNGE`, `UNSELECT`, `CLOSE`, `SUBSCRIBE`, `UNSUBSCRIBE` (including LIST's [RFC 9051](https://www.rfc-editor.org/rfc/rfc9051) section 6.3.9.1 `SUBSCRIBED` selection option), the `UID`-prefixed form of every command that supports it, `IDLE` (with the cross-session caveat noted below), and the [RFC 7162](https://www.rfc-editor.org/rfc/rfc7162) `CONDSTORE`/`QRESYNC` extensions.
ACL/shared-mailbox support is deliberately left out.
**`IDLE` is a poll, not a kernel-driven push.** An `IDLE`ing session rechecks its selected mailbox every `idle poll` seconds (default 5, see `imapd.conf`), so new mail, whether delivered by an external MTA or by another IMAP session, is reported within one interval rather than instantly. Most polls are two `stat(2)` calls and no lock: the store child only re-reads the index and streams UIDs when the mailbox directory or `new/` has actually been touched. Setting `idle poll 0` disables polling entirely, which restores the earlier behaviour where an `IDLE`ing session saw nothing until it sent `DONE`.
## Requirements
OpenBSD only. This depends on `<imsg.h>`, `pledge(2)`, `unveil(2)`, and libutil's `imsgbuf_*` API, none of which exist outside OpenBSD. Links against libevent, libtls/libssl/libcrypto, and libutil, all base-system libraries (see `src/Makefile`).
**imapd requires OpenBSD -current, and will not build on 7.9.**. Building on the most recent stable release is a goal for 1.0.
`keymgr`, the process that isolates the TLS private key from `listener`, additionally depends on `tls_config_use_fake_private_key()` and undocumented ex_data-tagging behavior inside `tls_keypair_load()`, both unexported libtls/LibreSSL internals with no compatibility promise, and neither declared in libtls's public, installed `tls.h`.
## Building and installing
```
cd src
make
doas make install
```
Installs the daemon to `/usr/local/sbin/imapd`, man pages to `/usr/local/man/man8`, the `imapduser` account-provisioning tool alongside the daemon, and a sample config to `/usr/local/share/examples/imapd/imapd.conf`.
The `rc.d(8)` script is not installed automatically. `install(1)`, not `cp(1)`, so the installed copy is executable regardless of the source tree's own permission bits:
```
doas install -o root -g wheel -m 555 src/rc.d/imapd /etc/rc.d/imapd
```
## Configuring
Copy the sample config into place with restrictive permissions. imapd refuses to start against a config that's group- or world-writable, *or* world-readable:
```
doas install -o root -g wheel -m 600 \
/usr/local/share/examples/imapd/imapd.conf /etc/imapd.conf
```
Every directive is documented inline in the sample file; the full reference is in `imapd(8)`'s FILES section.
## Creating an account
imapd's users aren't real system accounts, `imapduser(8)` manages a bespoke credentials file (`username:passwordhash:uid:gid:maildir`, bcrypt via `crypt_checkpass(3)`) and the matching maildir ownership together, since no combination of `useradd(8)`/`userdel(8)` can safely keep both in sync:
```
doas imapduser -a someuser
```
See `imapduser(8)` for `-d` (revoke login without touching mail) and the `-c`/`-s`/`-u`/`-g` overrides.
## Running
```
doas rcctl enable imapd
doas rcctl start imapd
```
## Known limitations
Beyond the deliberate protocol-scope decisions covered in `imapd(8)`'s CAVEATS:
- If the listener or auth process exits unexpectedly after startup, it is not automatically restarted. Recovery is `rcctl restart imapd`. See `imapd(8)`.
`SIGHUP` reloads `spool`, `append max`, `attachment max`, `account sessions`, `connections max`, `idle poll`, `lock timeout`, `login grace`, `startups`, and the TLS certificate/key without dropping connected sessions, `listen on` and `credentials` changes still require a restart. See `imapd(8)`.
IPv6 is supported (`listen on ::` or `listen on *` for dual-stack) but not the default, see `imapd(8)`'s `listen on` directive.
## Getting the source
The repository is hosted with [Game of Trees](https://gameoftrees.org/) (`got`), read-only anonymous access over SSH:
```
got clone ssh://anonymous@got.openimapd.dev/imapd
```
The repository is also git-compatible; a plain `git clone` against the same URL works too:
```
git clone ssh://anonymous@got.openimapd.dev/imapd
```
## Security
Report security issues to security@openimapd.dev. General questions or feedback: feedback@openimapd.dev.
## License
ISC. See the copyright header in each source file.
## More
`imapd(8)` and `imapduser(8)` are the authoritative technical reference.
