commit f940a630e00423a0043b84fbe5affd05ea548d6d from: David Williams date: Sun Sep 27 23:40:47 2026 UTC SEARCH BODY and TEXT, parse cache, RFC 5322 headers, and literals Ten changes, made one at a time and committed together. Each follows in the order it was made, under its own subject line. Together: the account worker keeps the ENVELOPE and BODYSTRUCTURE text it has had parsed; SEARCH answers BODY and TEXT, so it now answers every key of RFC 9051 section 6.4.4; ENVELOPE and SEARCH read addresses, groups and comments as RFC 5322 writes them; a string may be sent as a literal (RFC 9051 section 4.3) wherever imapd reads one, while a refused non-synchronizing literal's command now ends where RFC 7888 says it does; a header field is found when white space comes before its colon (RFC 5322 section 4.5); a bare CR in a quoted Content-Type parameter no longer loses that parameter and those after it; and SEARCH takes its CHARSET quoted, as RFC 9051 allows. Change 1 of 10: Cache ENVELOPE and BODYSTRUCTURE in the account worker Every ENVELOPE or BODYSTRUCTURE FETCH item cost a message open, a descriptor pass, a parse in the parser-worker and a wait for its reply, however often the same message had been asked for. On a mailbox of 10,000 messages, FETCH 1:* (ENVELOPE) took 4.4 times as long as FETCH 1:* (FLAGS), which needs no parser. The account worker now keeps the checked text of both items, in the shape of smtpd's envelope cache: a fixed bound with no directive, an entry moved to the front on each use, and the oldest evicted first. The bound is 4 MiB of text and bookkeeping per account worker. Only successes are kept; a message the parser could not answer is asked for again, and its strikes are counted as before. With the cache, a second FETCH 1:* (ENVELOPE) of the same 10,000 messages took about half as long as the first. The key is RFC 9051 section 2.3.1.1's: mailbox name, UIDVALIDITY and UID, which "must refer to a single, immutable (or expunged) message on that server forever". A hit must also match the message's file name in the index, so an index rewritten by another program, or a UIDVALIDITY issued twice, is a miss rather than another message's text. A hit still needs the message file on disk, as the uncached path does, so a FETCH answers as it did before for a message whose file has gone. EXPUNGE, CLOSE and MOVE drop the UIDs they end, DELETE and RENAME drop every entry of the name they end, and a FETCH drops its mailbox's entries when the first of them has another UIDVALIDITY. The key alone would make each of these cost memory rather than a wrong answer if missed. A message whose requested items need no parse, or are all in the cache, no longer waits for a parser-worker, so a warm FETCH of ENVELOPE or BODYSTRUCTURE does not start one. When an account worker that looked anything up exits, it logs one line with its hits, misses, additions, evictions, drops, stale entries found, and its peak entries and bytes, so the bound can be sized from what the cache holds. A worker that never fetched ENVELOPE or BODYSTRUCTURE logs nothing, so short sessions add no lines. Change 2 of 10: Answer SEARCH BODY and TEXT RFC 9051 section 6.4.4's BODY and TEXT were parsed and refused NO. They are now answered in the parser-worker, beside the header keys, from one request per candidate message as before; a request with either key reads the whole message, up to "attachment max", instead of the header alone. The message is walked as MIME. Only TEXT and MESSAGE parts are searched, as section 6.4.4 permits; base64 and quoted-printable are decoded first, as it requires. An unrecognized transfer encoding makes a part application/octet-stream (RFC 2045 section 6.4), so it is not searched. A missing or invalid Content-Type is text/plain (section 5.2), and so is a multipart with no boundary. An unrecognized multipart subtype is walked as mixed (RFC 2046 section 5.1.3), and a digest part with no Content-Type is MESSAGE/RFC822 (section 5.1.5). Preambles and epilogues are not searched. BODY matches no header field of any level: not the message's, not a MIME part's, not a forwarded message's. TEXT matches all of them, each field unfolded and its RFC 2047 words decoded, and the content BODY sees. A forwarded MESSAGE/RFC822 or MESSAGE/GLOBAL is walked into and its own parts decoded. HTML is searched as it is, markup included, and a phrase is matched only within one line of the decoded text. Message charsets are not converted; an ASCII word still matches beside other octets. A multipart has no limit on its parts; the message's size bounds them. An entity nested deeper than BODYSTRUCTURE's limit of 10 cannot be checked, so the SEARCH answers NO, as for any message the parser cannot check. Matching folds ASCII case and runs in time linear in the text (Knuth-Morris-Pratt), so what a crafted body and needle can cost grows with the text's length, not with the product of the two lengths. The message is read into a buffer that doubles rather than grows by 64K. Change 3 of 10: Parse RFC 5322 addresses for ENVELOPE and SEARCH as the RFC writes them ENVELOPE and SEARCH FROM, TO, CC and BCC share one address parse, and it did not know RFC 5322 section 3.4's groups or section 3.2.2's comments, and toggled on every double quote, escaped or not. So "team: a@b, c@d;" came out as a mailbox "team: a" and a host "d;"; an escaped quote in a display name dropped that address and every one after it; "jane@x.org (Doe, Jane)" split inside its comment; and a comment stayed in the host or the name, so SEARCH matched it. These were wrong answers with nothing to show for it. A group is now sent as RFC 9051 section 7.5.2 defines it: a marker holding the group's name before its members and an empty one after, an empty group included, so "undisclosed-recipients:;" is two markers where it was NIL. A group missing its ";" is closed at the end of the field. SEARCH still matches the members, not a group's name. Comments are removed before the list is split, each becoming one space, nested and with quoted-pairs, as mail(1)'s skin() drops them; a comment is never taken as a display name. A backslash escapes the next character in a quoted string. The display name loses its quoting and keeps single spaces between words; the local part loses its quoting; an address loses white space outside quotes, so the obsolete "jdoe@test . example" is "test.example". An obsolete route goes in the at-domain-list, not the mailbox, and SEARCH does not match it. An address that does not parse is still left out and the rest of the field kept. An address longer than 1024 octets is no longer dropped: only the whole ENVELOPE has a limit, as before. Change 4 of 10: End a refused non-synchronizing literal's command where RFC 7888 does A command whose non-synchronizing literal ("{n+}", RFC 9051 section 4.3) was refused, or answered before its end, had only the literal's octets discarded. The rest of the command, which RFC 9051 section 7.6 puts after the octets, was then read as a new command line, so "a SEARCH TEXT {5+}" followed by "hello UNSEEN" drew a second reply, "UNSEEN BAD Missing command", tagged with the client's own word. RFC 7888 section 3 requires the octets and the following line to be treated as part of the same command. The line after the octets is now discarded too, and if it ends in another "{n+}" those octets are discarded in turn. A non-synchronizing literal over 4096 octets still closes the connection, one of the two choices RFC 7888 section 4 allows, but the reply is now "* BYE [TOOBIG]" as that choice and section 5 describe, where it was "* BAD". APPEND's own refusal of a non-synchronizing literal over 4096 octets is removed. The listener closes the connection on such a line before any command sees it, so the check could not be reached. Change 5 of 10: Read literals as mailbox names and SEARCH strings RFC 9051 section 4.3 makes a literal, "{n}" or "{n+}" followed by n octets, one of the two forms of a string, and a client may send one wherever a mailbox name or a SEARCH string may go; the RFC's own SEARCH examples in section 6.4.4 do. imapd took a literal only as APPEND's message and answered BAD to every other, so a client that chose the literal form could not select, create or search. A parser that meets a literal ending the text it was given now asks for it instead of refusing it. The listener then gathers the command: it sends "+" for a synchronizing literal, reads the n octets and the line after them, and runs the whole command again from the start, as ldapd does with a request whose BER element is not all there yet. The parsers read a literal already in the gathered text by its count, so its octets may hold a space, a quote or a CRLF. Each command parses all of its arguments before it acts, so asking for a literal changes nothing; SEARCH now parses before it resets the session's last result. A command and its literals are held to 8191 octets. Over that, it is answered NO [LIMIT], before the "+" when the count is known in time. A NUL in a literal, which CHAR8 excludes, is answered BAD. In either case the rest of the command is discarded as RFC 7888 section 3 asks. This covers every mailbox argument, LIST's reference and pattern, SEARCH's strings and HEADER's field name. FETCH's HEADER.FIELDS names and ID's parameters still do not take a literal. Change 6 of 10: Take FETCH's header field names quoted or as literals, and gather ID's A HEADER.FIELDS name is an astring (RFC 9051 section 9, header-fld-name), but FETCH took it as an atom only: a quoted name was BAD, and a literal never reached the parser. Names are now read quoted, with RFC 9051's two escapes, or as literals, which FETCH turns into quoted strings before it parses, since a field name cannot hold the CR or LF only a literal could carry. The response echoes the list as the client sent it, a literal shown as the quoted string it became. A name that would hold a space is BAD, as the store takes the list space-separated, and so is an empty list, which header-list does not allow. The FETCH tokenizer no longer splits inside a quoted string. ID answered OK without reading its parameters, so a client that sent one as a synchronizing literal got the reply before its "+". It now reads just far enough to gather its literals (RFC 2971 section 3.1) and still answers NIL. Change 7 of 10: Accept a non-synchronizing literal on a pipelined command A command that arrived while another was in flight, and that carried a non-synchronizing literal, was refused with an untagged BAD, so its tag was never answered; RFC 9051 section 2.2.2 tags every completion, and section 4.3 lets a client send "{n+}" anywhere a literal may go. Its octets and the lines after them are now gathered into the queued command, held to the same 8191 octets as any other, and it runs in its turn like any pipelined command (section 5.5). A queued command also no longer waits for the store's next reply when the gathered command ahead of it finished without needing the store. Change 8 of 10: Find a header field whose name is followed by white space RFC 5322 section 4.5 lets any field put white space between its name and its colon, "From :" for "From:", and section 4 requires a receiver to accept that. imapd took a field's name to be everything before the colon, white space included, so such a field was never found: ENVELOPE showed NIL, SEARCH did not match it, HEADER.FIELDS left it out, and a "Content-Type :" multipart message was described and searched as one text/plain part. The name now ends before any SP or HTAB that precedes the colon, in the one function every header reader uses. HEADER.FIELDS still returns the field as the message holds it. Change 9 of 10: Keep a Content-Type parameter that holds a bare CR A CR, LF or NUL inside a quoted Content-Type parameter value made the quoted-string fail, and the parameter loop stopped there, so that parameter and every one after it were lost. text/plain; charset="usascii" was shown with NIL parameters, and a multipart whose boundary came after such a parameter, or held the CR itself, could not be split: its BODYSTRUCTURE ended NO, its BODY[n] was empty, and SEARCH read the whole body as text, delimiter lines included. RFC 5322 section 4 says malformed data is not to be "irretrievably lost". The value is now kept, CR and LF included, so a boundary still matches delimiter lines that carry the same CR. A NUL becomes a space, since the value is held as a C string. Every parameter value already goes out through the nstring writer the other fields use, which sends a CR, LF or NUL as a space, since a quoted string cannot hold one (RFC 9051 section 9). A message holding a NUL is still refused for BODYSTRUCTURE, BODY[] and BODY[n], so a NUL in a parameter now matters only to SEARCH. Change 10 of 10: Read SEARCH's CHARSET as an atom or a quoted string RFC 9051 section 9 writes SEARCH's charset as "atom / quoted", but imapd took every octet up to the next space. So CHARSET "UTF-8" was compared with its quotes and refused NO [BADCHARSET], and a name the grammar does not allow, such as UTF(8 or a quoted string with no closing quote, was answered NO where its arguments are invalid, which section 6.4.4 answers BAD. The charset is now read as the grammar writes it: a quoted string, whose only escapes are \" and \\, or an atom, which refuses the atom-specials, controls and 8-bit octets; a space or the end of the command must follow it. A malformed charset is BAD. A well-formed one other than US-ASCII or UTF-8 is still NO [BADCHARSET], as section 6.4.4 requires. commit - 7d9334a7d40e0d8a8bbba80f1e1026a34ec35911 commit + f940a630e00423a0043b84fbe5affd05ea548d6d blob - 0ecebc52b885ae539520e05c47e61857f7d8be59 blob + 8c9c9862d382bfe7df46a4cc864956499b97ce3a --- README.md +++ README.md @@ -2,7 +2,7 @@ A from-scratch IMAP4rev2 ([RFC 9051](https://www.rfc-editor.org/rfc/rfc9051)) server for OpenBSD, written in C in the privilege-separated tradition of `smtpd(8)`, `httpd(8)`, and `ntpd(8)`. No third-party IMAP library. -**Status:** Pre-release, actively developed. Not a port. See [Getting the source](#getting-the-source) below for the repository. +**Status:** 0.1.7 Pre-release, actively developed. Not a port. See [Getting the source](#getting-the-source) below for the repository. ## What it is blob - 275b0f615d50095ed9373998a6bc4fa669c369b6 blob + 60bed43482a76e1603b8d438342657969d72b840 --- contrib/imapduser.8 +++ contrib/imapduser.8 @@ -2,7 +2,7 @@ .\" .\" Written for the OpenIMAPD project. Public domain / no rights reserved. .\" -.Dd $Mdocdate: September 25 2026 $ +.Dd $Mdocdate: September 27 2026 $ .Dt IMAPDUSER 8 .Os .Sh NAME blob - d5bf488c9c64067329371bc773787b88e8e5e738 blob + 4bf7b96ec5aa1ffb1dbf3308e5cf4a0445e3b7a0 --- src/Makefile +++ src/Makefile @@ -13,7 +13,8 @@ SRCS= main.c parent.c log.c imsgev.c parse.y utf8.c m keymgr.c \ parser.c \ store.c index.c mime.c envelope.c mbox_fetch.c mbox_search.c \ - mbox_store.c mbox_manage.c mbox_copy.c search_match.c + mbox_store.c mbox_manage.c mbox_copy.c search_match.c \ + store_cache.c BINDIR= /usr/local/sbin MANDIR= /usr/local/man/man blob - b6d33d4ccde5be554ed9a3d89cfe51ff2455903e blob + f7a13b2a677ece560279a205ca1d8a83d19722fd --- src/append_cmd.c +++ src/append_cmd.c @@ -221,14 +221,6 @@ parse_append_args(char *args, struct append_parsed *ou } out->litlen = (uint64_t)litlen; - if (out->litnonsync && out->litlen > 4096) { - /* RFC 9051 SS4.3 caps a non-synchronizing literal */ - *errmsg = "non-synchronizing literal exceeds RFC " - "9051 SS4.3's 4096-octet limit, use a " - "synchronizing literal instead"; - return (-1); - } - if (end[1] != '\0') { *errmsg = "literal must be the final argument"; return (-1); @@ -247,10 +239,8 @@ cmd_append(struct session *s, const char *tag, char *a const char *errmsg; rc = parse_append_args(args, &parsed, &errmsg); - if (rc == -1) { - session_reply(s, tag, "BAD", errmsg); - return (1); - } + if (rc == -1) + return (session_arg_error(s, tag, errmsg)); if (rc == -2) { session_reply(s, tag, "NO", errmsg); return (1); blob - d12b562475f766678ff91628fbb404cb544a5d9c blob + 9bdf08c3065723587c43f234b227be200861d536 --- src/auth_cmd.c +++ src/auth_cmd.c @@ -85,7 +85,24 @@ cmd_logout(struct session *s, const char *tag, char *a int cmd_id(struct session *s, const char *tag, char *args) { - /* RFC 2971 SS3.1: logged, not parsed; the reply is NIL */ + char *p = args, *lit; + size_t len; + const char *errmsg; + int inquote = 0; + + /* RFC 2971 SS3.1: only literals are read; the reply is NIL */ + while (p != NULL && *p != '\0') { + if (!inquote && *p == '{') { + if (read_literal(&p, &lit, &len, &errmsg) == -1) + return (session_arg_error(s, tag, errmsg)); + continue; + } + if (inquote && *p == '\\' && p[1] != '\0') + p++; + else if (*p == '"') + inquote = !inquote; + p++; + } log_debug("session %u: ID params: %s", s->id, args != NULL ? args : "(none)"); session_untagged(s, "ID NIL"); blob - 724a37ec1dd67f91057de94d728c98a17534e5f5 blob + a5728b7e6970463d7e16080c594c7fb8719cb862 --- src/envelope.c +++ src/envelope.c @@ -37,6 +37,9 @@ #include "log.h" #include "store_internal.h" +static size_t quoted_end(const char *, size_t, size_t); +static size_t unquoted_find(const char *, size_t, size_t, const char *); + int envbuf_append(char *buf, size_t bufsize, size_t *outlen, const char *data, size_t datalen) @@ -91,280 +94,228 @@ fail: return (-1); } -int -address_split(const char *tok, size_t toklen, const char **name_out, - size_t *namelen_out, const char **mailbox_out, size_t *mailboxlen_out, - const char **host_out, size_t *hostlen_out) +/* RFC 5322 SS3.2.4: the index just past the quoted string at i */ +static size_t +quoted_end(const char *s, size_t len, size_t i) { - const char *name = NULL; - size_t namelen = 0; - const char *spec; - size_t speclen; - const char *mailbox, *host; - size_t mailboxlen, hostlen; - size_t at; - int found_at; - size_t lt; + for (i++; i < len; i++) { + if (s[i] == '\\' && i + 1 < len) + i++; + else if (s[i] == '"') + return (i + 1); + } + return (len); +} - while (toklen > 0 && (tok[0] == ' ' || tok[0] == '\t')) { - tok++; - toklen--; +static size_t +unquoted_find(const char *s, size_t len, size_t i, const char *chars) +{ + while (i < len) { + if (s[i] == '"') + i = quoted_end(s, len, i); + else if (s[i] != '\0' && strchr(chars, s[i]) != NULL) + return (i); + else + i++; } - while (toklen > 0 && (tok[toklen - 1] == ' ' || - tok[toklen - 1] == '\t')) - toklen--; - if (toklen == 0) - return (-1); + return (len); +} - /* unquoted '<' splits display-name from addr-spec (to unquoted '>') */ - lt = toklen; - { - size_t i; - int q = 0; +/* RFC 5322 SS3.2.2: each comment becomes one space; returns the length */ +size_t +address_uncomment(char *s, size_t len) +{ + size_t i = 0, n = 0, e; + int depth; - for (i = 0; i < toklen; i++) { - if (tok[i] == '"') - q = !q; - else if (!q && tok[i] == '<') { - lt = i; - break; + while (i < len) { + if (s[i] == '"') { + e = quoted_end(s, len, i); + memmove(s + n, s + i, e - i); + n += e - i; + i = e; + } else if (s[i] == '(') { + for (depth = 0; i < len; i++) { + if (s[i] == '\\' && i + 1 < len) + i++; + else if (s[i] == '(') + depth++; + else if (s[i] == ')' && --depth == 0) { + i++; + break; + } } + s[n++] = ' '; + } else + s[n++] = s[i++]; + } + return (n); +} + +/* unquoted in place (RFC 5322 SS3.2.4); a phrase keeps one space per run */ +size_t +address_cook(char *s, size_t len, int phrase) +{ + size_t i, e, n = 0; + int space = 0; + + for (i = 0; i < len; i++) { + if (s[i] == ' ' || s[i] == '\t') { + space = phrase; + continue; } + if (space && n > 0) + s[n++] = ' '; + space = 0; + if (s[i] != '"') { + s[n++] = s[i]; + continue; + } + for (e = quoted_end(s, len, i), i++; i < e; i++) { + if (s[i] == '\\' && i + 1 < e) + i++; + else if (s[i] == '"') + break; + s[n++] = s[i]; + } } + return (n); +} +/* RFC 5322 SS3.4 mailbox, split and cooked in place; -1 if it is none */ +int +address_split(char *tok, size_t toklen, struct address *a) +{ + char *spec = tok; + size_t speclen = toklen, lt, gt, colon, at, i; + + memset(a, 0, sizeof(*a)); + lt = unquoted_find(tok, toklen, 0, "<"); if (lt < toklen) { - size_t gt = toklen, i; - int q = 0; - - for (i = lt + 1; i < toklen; i++) { - if (tok[i] == '"') - q = !q; - else if (!q && tok[i] == '>') { - gt = i; - break; - } - } - if (gt >= toklen) - return (-1); /* unmatched '<', malformed, skip */ - - { - const char *disp = tok; - size_t displen = lt; - - while (displen > 0 && - (disp[0] == ' ' || disp[0] == '\t')) { - disp++; - displen--; - } - while (displen > 0 && - (disp[displen - 1] == ' ' || - disp[displen - 1] == '\t')) - displen--; - - if (displen >= 2 && disp[0] == '"' && - disp[displen - 1] == '"') { - disp++; - displen -= 2; - } - if (displen > 0) { - name = disp; - namelen = displen; - } - } - + gt = unquoted_find(tok, toklen, lt + 1, ">"); + if (gt == toklen) + return (-1); spec = tok + lt + 1; - speclen = gt - (lt + 1); - } else { - spec = tok; - speclen = toklen; + speclen = gt - lt - 1; + if ((a->namelen = address_cook(tok, lt, 1)) > 0) + a->name = tok; } - while (speclen > 0 && (spec[0] == ' ' || spec[0] == '\t')) { spec++; speclen--; } - while (speclen > 0 && (spec[speclen - 1] == ' ' || - spec[speclen - 1] == '\t')) - speclen--; - found_at = 0; - at = 0; - { - size_t i; - int q = 0; - - for (i = 0; i < speclen; i++) { - if (spec[i] == '"') - q = !q; - else if (!q && spec[i] == '@') { - at = i; - found_at = 1; - } - } + /* RFC 5322 SS4.4 obs-route, for RFC 9051's at-domain-list */ + if (speclen > 0 && spec[0] == '@') { + colon = unquoted_find(spec, speclen, 0, ":"); + if (colon == speclen) + return (-1); + a->route = spec; + a->routelen = address_cook(spec, colon, 0); + spec += colon + 1; + speclen -= colon + 1; } - if (!found_at || at == 0 || at + 1 >= speclen) - return (-1); /* no usable local-part@domain split */ - mailbox = spec; - mailboxlen = at; - host = spec + at + 1; - hostlen = speclen - at - 1; - - /* unquote a quoted local-part */ - if (mailboxlen >= 2 && mailbox[0] == '"' && - mailbox[mailboxlen - 1] == '"') { - mailbox++; - mailboxlen -= 2; - } - - *name_out = name; - *namelen_out = namelen; - *mailbox_out = mailbox; - *mailboxlen_out = mailboxlen; - *host_out = host; - *hostlen_out = hostlen; + if (unquoted_find(spec, speclen, 0, ":;<>") < speclen) + return (-1); + at = speclen; + for (i = unquoted_find(spec, speclen, 0, "@"); i < speclen; + i = unquoted_find(spec, speclen, i + 1, "@")) + at = i; + if (at == speclen || at == 0) + return (-1); + a->host = spec + at + 1; + if ((a->hostlen = address_cook(a->host, speclen - at - 1, 0)) == 0) + return (-1); + a->mailbox = spec; + a->mailboxlen = address_cook(spec, at, 0); return (0); } -/* RFC 9051 SS9 address; no groups */ +/* RFC 9051 SS7.5.2 address structure; a group marker has no host */ int envbuf_append_one_address(char *buf, size_t bufsize, size_t *outlen, - const char *tok, size_t toklen) + const struct address *a) { - char addrbuf[1024]; - size_t addrlen = 0; - const char *name, *mailbox, *host; - size_t namelen, mailboxlen, hostlen; - - if (address_split(tok, toklen, &name, &namelen, &mailbox, &mailboxlen, - &host, &hostlen) == -1) + if (envbuf_append(buf, bufsize, outlen, "(", 1) == -1 || + envbuf_append_nstring(buf, bufsize, outlen, a->name, + a->namelen) == -1 || + envbuf_append(buf, bufsize, outlen, " ", 1) == -1 || + envbuf_append_nstring(buf, bufsize, outlen, a->route, + a->routelen) == -1 || + envbuf_append(buf, bufsize, outlen, " ", 1) == -1 || + envbuf_append_nstring(buf, bufsize, outlen, a->mailbox, + a->mailboxlen) == -1 || + envbuf_append(buf, bufsize, outlen, " ", 1) == -1 || + envbuf_append_nstring(buf, bufsize, outlen, a->host, + a->hostlen) == -1) return (-1); - - if (envbuf_append(addrbuf, sizeof(addrbuf), &addrlen, "(", 1) == -1) - return (-1); - - if (name != NULL) { - size_t j; - - if (envbuf_append(addrbuf, sizeof(addrbuf), &addrlen, - "\"", 1) == -1) - return (-1); - for (j = 0; j < namelen; j++) { - char c = name[j]; - - if (c == '\\' && j + 1 < namelen) { - j++; - c = name[j]; - } - /* substituted as in envbuf_append_nstring() */ - if (c == '\0' || c == '\r' || c == '\n') - c = ' '; - if ((c == '"' || c == '\\') && - envbuf_append(addrbuf, sizeof(addrbuf), &addrlen, - "\\", 1) == -1) - return (-1); - if (envbuf_append(addrbuf, sizeof(addrbuf), &addrlen, - &c, 1) == -1) - return (-1); - } - if (envbuf_append(addrbuf, sizeof(addrbuf), &addrlen, - "\"", 1) == -1) - return (-1); - } else { - if (envbuf_append_str(addrbuf, sizeof(addrbuf), &addrlen, - "NIL") == -1) - return (-1); - } - - if (envbuf_append_str(addrbuf, sizeof(addrbuf), &addrlen, - " NIL ") == -1) - return (-1); - if (envbuf_append_nstring(addrbuf, sizeof(addrbuf), &addrlen, mailbox, - mailboxlen) == -1) - return (-1); - if (envbuf_append(addrbuf, sizeof(addrbuf), &addrlen, " ", 1) == -1) - return (-1); - if (envbuf_append_nstring(addrbuf, sizeof(addrbuf), &addrlen, host, - hostlen) == -1) - return (-1); - if (envbuf_append(addrbuf, sizeof(addrbuf), &addrlen, ")", 1) == -1) - return (-1); - - return (envbuf_append(buf, bufsize, outlen, addrbuf, addrlen)); + return (envbuf_append(buf, bufsize, outlen, ")", 1)); } +/* RFC 5322 SS3.4 address-list: the next mailbox, group start or end */ int -address_list_next(const char *val, size_t vallen, size_t *pos, - const char **tok_out, size_t *toklen_out) +address_list_next(char *val, size_t vallen, size_t *pos, int *ingroup, + char **tok_out, size_t *toklen_out) { - size_t i = *pos; + size_t i = *pos, start; - while (i < vallen) { - size_t tok_start; - size_t tok_len; - int in_quotes = 0, in_angle = 0; + while (i < vallen && (val[i] == ' ' || val[i] == '\t' || + val[i] == ',')) + i++; + if (i == vallen || (*ingroup && val[i] == ';')) { + *pos = i < vallen ? i + 1 : i; + if (*ingroup == 0) + return (0); + *ingroup = 0; + return (ADDR_GROUP_END); + } - while (i < vallen && (val[i] == ' ' || val[i] == '\t' || - val[i] == ',')) - i++; - tok_start = i; - while (i < vallen) { - char c = val[i]; + start = i; + while ((i = unquoted_find(val, vallen, i, *ingroup ? "<,;" : + "<,:")) < vallen && val[i] == '<') + i = unquoted_find(val, vallen, i + 1, ">"); - if (c == '"') - in_quotes = !in_quotes; - else if (!in_quotes && c == '<') - in_angle = 1; - else if (!in_quotes && c == '>') - in_angle = 0; - else if (!in_quotes && !in_angle && c == ',') - break; - i++; - } - tok_len = i - tok_start; - while (tok_len > 0 && (val[tok_start + tok_len - 1] == ' ' || - val[tok_start + tok_len - 1] == '\t')) - tok_len--; - - if (tok_len > 0) { - *tok_out = val + tok_start; - *toklen_out = tok_len; - *pos = i; - return (1); - } + *tok_out = val + start; + *toklen_out = i - start; + while (*toklen_out > 0 && (val[start + *toklen_out - 1] == ' ' || + val[start + *toklen_out - 1] == '\t')) + (*toklen_out)--; + if (i < vallen && val[i] == ':') { + *pos = i + 1; + *ingroup = 1; + return (ADDR_GROUP); } - *pos = i; - return (0); + *pos = i < vallen && val[i] == ',' ? i + 1 : i; + return (ADDR_MAILBOX); } -/* RFC 9051 SS7.5.2 address list, or NIL */ +/* RFC 9051 SS7.5.2 address list, or NIL; val is rewritten in place */ int envbuf_append_address_list(char *buf, size_t bufsize, size_t *outlen, - const char *val, size_t vallen) + char *val, size_t vallen) { - const char *tok; - size_t save = *outlen; - size_t i = 0, toklen; - int any = 0; + struct address a; + char *tok; + size_t save = *outlen, pos = 0, toklen; + int any = 0, ingroup = 0, kind; - while (vallen > 0 && (val[0] == ' ' || val[0] == '\t')) { - val++; - vallen--; - } - while (vallen > 0 && (val[vallen - 1] == ' ' || - val[vallen - 1] == '\t')) - vallen--; - - if (vallen == 0) - return (envbuf_append_str(buf, bufsize, outlen, "NIL")); - + vallen = address_uncomment(val, vallen); if (envbuf_append(buf, bufsize, outlen, "(", 1) == -1) return (-1); - while (address_list_next(val, vallen, &i, &tok, &toklen) == 1) { - if (envbuf_append_one_address(buf, bufsize, outlen, tok, - toklen) == 0) - any = 1; + while ((kind = address_list_next(val, vallen, &pos, &ingroup, &tok, + &toklen)) != 0) { + memset(&a, 0, sizeof(a)); + if (kind == ADDR_GROUP) { + a.mailbox = tok; + a.mailboxlen = address_cook(tok, toklen, 1); + } else if (kind == ADDR_MAILBOX && + address_split(tok, toklen, &a) == -1) + continue; + if (envbuf_append_one_address(buf, bufsize, outlen, &a) == -1) + return (-1); + any = 1; } if (!any) { @@ -423,7 +374,6 @@ build_envelope(int fd, const char *basename, char **bu if (envbuf_append(out, sizeof(out), &outlen, " ", 1) == -1) goto fail; - /* from */ { char *val; size_t vallen; @@ -447,7 +397,6 @@ build_envelope(int fd, const char *basename, char **bu if (envbuf_append(out, sizeof(out), &outlen, " ", 1) == -1) goto fail; - /* sender, reply-to: default to from_formatted per SS7.5.2 */ { static const char *const fallback_fields[] = { "Sender", "Reply-To" }; @@ -481,7 +430,6 @@ build_envelope(int fd, const char *basename, char **bu } } - /* to, cc, bcc */ { static const char *const addr_fields[] = { "To", "Cc", "Bcc" }; size_t fi; blob - 024c85828cc7b50984c280f0be02c9abad1a68af blob + fa0815a7ea209ae6c3da6c3e18505aba39d017ca --- src/fetch_cmd.c +++ src/fetch_cmd.c @@ -149,7 +149,7 @@ parse_sequence_set(const char *text, struct seq_range return (0); } -/* a space inside an unclosed "[" or "(" is not a delimiter */ +/* a space inside "[", "(" or a quoted string is not a delimiter */ static char * fetch_att_tok(char *str, char **savep) { @@ -166,7 +166,13 @@ fetch_att_tok(char *str, char **savep) } for (start = p; *p != '\0'; p++) { - if (*p == '[' || *p == '(') + if (*p == '"') { + while (*++p != '\0' && *p != '"') + if (*p == '\\' && p[1] != '\0') + p++; + if (*p == '\0') + break; + } else if (*p == '[' || *p == '(') depth++; else if (*p == ']' || *p == ')') { if (depth > 0) @@ -188,46 +194,52 @@ int parse_header_fields_att(const char *inner, int *not_out, char *fields_out, size_t fields_outsize) { - char listbuf[HEADER_FIELDS_LABEL_MAX]; - char *p = listbuf, *listp, *end; - char *name, *save; - int first = 1; + char name[HEADER_FIELDS_MAX]; + const char *p = inner; + size_t n; + int first = 1; *not_out = 0; fields_out[0] = '\0'; - if (strlcpy(listbuf, inner, sizeof(listbuf)) >= sizeof(listbuf)) + if (strlen(inner) >= HEADER_FIELDS_LABEL_MAX) return (-1); - if (strncasecmp(p, "HEADER.FIELDS", 13) != 0) return (-1); p += 13; - if (strncasecmp(p, ".NOT", 4) == 0) { *not_out = 1; p += 4; } - - if (*p != ' ') + if (p[0] != ' ' || p[1] != '(') return (-1); - p++; + p += 2; - listp = p; - if (*listp != '(') - return (-1); - listp++; - - end = strchr(listp, ')'); - if (end == NULL || end[1] != '\0') - return (-1); - *end = '\0'; - - if (*listp == '\0') - return (-1); /* header-list requires at least one name */ - - for (name = strtok_r(listp, " ", &save); name != NULL; - name = strtok_r(NULL, " ", &save)) { - if (strchr(name, '"') != NULL) + for (;;) { + while (*p == ' ') + p++; + if (*p == ')' || *p == '\0') + break; + n = 0; + if (*p == '"') { + for (p++; *p != '"'; p++) { + if (*p == '\0' || n + 1 >= sizeof(name)) + return (-1); + if (*p == '\\' && *++p != '"' && *p != '\\') + return (-1); + name[n++] = *p; + } + p++; + } else { + for (; *p != '\0' && *p != ' ' && *p != ')'; p++) { + if (*p == '"' || n + 1 >= sizeof(name)) + return (-1); + name[n++] = *p; + } + } + name[n] = '\0'; + /* the store takes the names space-separated */ + if (n == 0 || strchr(name, ' ') != NULL) return (-1); if (!first && strlcat(fields_out, " ", fields_outsize) >= fields_outsize) @@ -236,7 +248,8 @@ parse_header_fields_att(const char *inner, int *not_ou return (-1); first = 0; } - + if (first || *p != ')' || p[1] != '\0') + return (-1); return (0); } @@ -897,6 +910,49 @@ parse_fetch_modifiers(char *modspec, struct imsg_mbox_ return (0); } +/* RFC 9051 SS4.3: FETCH's literals, header-fld-names, as quoted strings */ +static int +fetch_unliteral(char *spec, char *out, size_t outsize, const char **errmsg) +{ + char *p = spec, *lit; + size_t o = 0, len, i; + int inquote = 0; + + while (*p != '\0') { + if (!inquote && *p == '{') { + if (read_literal(&p, &lit, &len, errmsg) == -1) + return (-1); + if (o + 2 * len + 2 >= outsize) { + *errmsg = "FETCH arguments too long"; + return (-1); + } + out[o++] = '"'; + for (i = 0; i < len; i++) { + if (lit[i] == '\r' || lit[i] == '\n') { + *errmsg = "CR or LF in a field name"; + return (-1); + } + if (lit[i] == '"' || lit[i] == '\\') + out[o++] = '\\'; + out[o++] = lit[i]; + } + out[o++] = '"'; + continue; + } + if (o + 2 >= outsize) { + *errmsg = "FETCH arguments too long"; + return (-1); + } + if (inquote && *p == '\\' && p[1] != '\0') + out[o++] = *p++; + else if (*p == '"') + inquote = !inquote; + out[o++] = *p++; + } + out[o] = '\0'; + return (0); +} + int cmd_fetch(struct session *s, const char *tag, char *args) { @@ -916,6 +972,7 @@ fetch_dispatch(struct session *s, const char *tag, cha int header_fields_not = 0; char header_fields[HEADER_FIELDS_MAX]; char header_fields_label[HEADER_FIELDS_LABEL_MAX]; + char unlit[SESSION_INBUF_MAX]; int bodystructure_full = 0; char section_part[SECTION_PART_MAX]; int has_partial = 0; @@ -946,6 +1003,9 @@ fetch_dispatch(struct session *s, const char *tag, cha session_reply(s, tag, "BAD", errmsg); return (1); } + if (fetch_unliteral(attspec, unlit, sizeof(unlit), &errmsg) == -1) + return (session_arg_error(s, tag, errmsg)); + attspec = unlit; modspec = split_trailing_modifiers(attspec); blob - fe08b5b977abd9e830855e87f2de1c8f2f51072f blob + e17940a4005eba802bafef76c6836b6948719c4f --- src/imapd.8 +++ src/imapd.8 @@ -3,7 +3,7 @@ .\" Written for the OpenIMAPD project. Public domain / no rights reserved, .\" matching the project's ports-oriented, OpenBSD-base-inclusion goal. .\" -.Dd $Mdocdate: September 25 2026 $ +.Dd $Mdocdate: September 27 2026 $ .Dt IMAPD 8 .Os .Sh NAME blob - 36a55e455f441bce94458123f671b274b6e98577 blob + 8b8d979bbea17a86332be12e7642d65a0584705f --- src/imapd.h +++ src/imapd.h @@ -28,7 +28,7 @@ #include #include -#define IMAPD_VERSION "0.1.6" +#define IMAPD_VERSION "0.1.7" enum openimap_proc_type { PROC_PARENT, @@ -620,6 +620,8 @@ struct imsg_mbox_appended { #define SEARCH_OP_TO 30 #define SEARCH_OP_CC 31 #define SEARCH_OP_BCC 32 +#define SEARCH_OP_BODY 33 +#define SEARCH_OP_TEXT 34 struct search_node { int op; blob - a44a4d1a1971a77be423a9b5bb97ba7fe47a958c blob + b83f6daf82d8b6e6b1b0d41ecc8fb9295a45c899 --- src/listener.c +++ src/listener.c @@ -826,37 +826,175 @@ static int session_enqueue_cmd(struct session *, const /* RFC 9051 SS4.3 hard cap on a non-synchronizing literal. */ #define IMAP_NONSYNC_LITERAL_MAX 4096 +/* RFC 9051 SS4.3 "{n}" or "{n+}" at p; *endp is past its "}" */ static int -line_nonsync_literal(const char *line, uint64_t *lenp) +literal_count(const char *p, uint64_t *lenp, int *nonsyncp, const char **endp) { - const char *open, *stop; - char digits[24], *end; - size_t len, dlen; + char digits[24]; + size_t dlen; unsigned long long v; - *lenp = 0; - len = strlen(line); - if (len < 4 || line[len - 1] != '}' || line[len - 2] != '+') + if (*p++ != '{') return (0); - stop = &line[len - 2]; - if ((open = memrchr(line, '{', len)) == NULL || open + 1 >= stop) + dlen = strspn(p, "0123456789"); + if (dlen == 0 || dlen >= sizeof(digits)) return (0); - open++; - dlen = (size_t)(stop - open); - if (dlen >= sizeof(digits) || *open < '0' || *open > '9') - /* also rejects strtoull(3)'s sign/space forms */ - return (0); - memcpy(digits, open, dlen); + memcpy(digits, p, dlen); digits[dlen] = '\0'; - errno = 0; - v = strtoull(digits, &end, 10); - if (*end != '\0' || errno == ERANGE) + v = strtoull(digits, NULL, 10); + if (errno == ERANGE) return (0); + p += dlen; + *nonsyncp = (*p == '+'); + if (*nonsyncp) + p++; + if (*p != '}') + return (0); *lenp = (uint64_t)v; + *endp = p + 1; return (1); } +static int +line_literal(const char *line, uint64_t *lenp, int *nonsyncp) +{ + const char *open, *end; + size_t len; + + len = strlen(line); + if (len < 3 || line[len - 1] != '}' || + (open = memrchr(line, '{', len)) == NULL) + return (0); + return (literal_count(open, lenp, nonsyncp, &end) && *end == '\0'); +} + +static int +line_nonsync_literal(const char *line, uint64_t *lenp) +{ + int nonsync; + + if (!line_literal(line, lenp, &nonsync) || !nonsync) { + *lenp = 0; + return (0); + } + return (1); +} + +const char literal_wanted[] = "literal wanted"; + +/* a literal inside a gathered command, or ask for it if it ends the text */ +int +read_literal(char **pp, char **datap, size_t *lenp, const char **errmsg) +{ + const char *end; + uint64_t n; + int nonsync; + + if (!literal_count(*pp, &n, &nonsync, &end)) { + *errmsg = "malformed literal"; + return (-1); + } + if (*end == '\0') { + *errmsg = literal_wanted; + return (-1); + } + if (end[0] != '\r' || end[1] != '\n' || + (uint64_t)strlen(end + 2) < n) { + *errmsg = "malformed literal"; + return (-1); + } + *datap = *pp + (end + 2 - *pp); + *lenp = (size_t)n; + *pp = *datap + *lenp; + return (0); +} + +/* BAD for a malformed argument, or wait for the literal it ends in */ +int +session_arg_error(struct session *s, const char *tag, const char *errmsg) +{ + if (errmsg == literal_wanted) + s->literal_wanted = 1; + else + session_reply(s, tag, "BAD", errmsg); + return (1); +} + +static const char cmd_too_long[] = "[LIMIT] command with its literals " + "exceeds 8191 octets"; + +static void +session_cmd_reply(struct session *s, const char *status, const char *text) +{ + char tag[IMAP_TAG_MAX]; + size_t skip, len; + + skip = strspn(s->cmdbuf, " "); + len = strcspn(s->cmdbuf + skip, " "); + if (len >= sizeof(tag)) + len = sizeof(tag) - 1; + memcpy(tag, s->cmdbuf + skip, len); + tag[len] = '\0'; + session_reply(s, tag, status, text); +} + +/* append the next literal's octets to cmdbuf, or refuse the command */ +static int +session_gather(struct session *s, uint64_t n) +{ + if (s->cmdlen + 3 > sizeof(s->cmdbuf) || + n > sizeof(s->cmdbuf) - 3 - s->cmdlen) { + session_cmd_reply(s, "NO", cmd_too_long); + s->cmd_queued = 0; + return (0); + } + memcpy(s->cmdbuf + s->cmdlen, "\r\n", 2); + s->cmdlen += 2; + s->cmd_octets = (size_t)n; + s->cmd_gather = 1; + return (1); +} + +/* RFC 9051 SS4.3: run a command, re-parsed as each literal arrives */ +static int +session_run_cmd(struct session *s, char *cmd) +{ + static const char cont[] = "+ Ready for literal data\r\n"; + static char cmdwork[SESSION_INBUF_MAX]; + uint64_t n; + int nonsync, alive; + + if (!line_literal(cmd, &n, &nonsync)) + return (session_handle_line(s, cmd)); + if (cmd != s->cmdbuf) + (void)strlcpy(s->cmdbuf, cmd, sizeof(s->cmdbuf)); + (void)strlcpy(cmdwork, s->cmdbuf, sizeof(cmdwork)); + s->literal_wanted = 0; + alive = session_handle_line(s, cmdwork); + if (!alive || !s->literal_wanted) + return (alive); + s->literal_wanted = 0; + s->cmdlen = strlen(s->cmdbuf); + if (session_gather(s, n) && !nonsync) + session_write(s, cont, sizeof(cont) - 1); + return (1); +} + +static int +session_queue(struct session *s, const char *cmd) +{ + static const char bad[] = "* BAD too many pipelined commands, " + "closing connection\r\n"; + + if (session_enqueue_cmd(s, cmd)) + return (1); + log_warnx("session %u: pipelined command queue full, closing", s->id); + session_write(s, bad, sizeof(bad) - 1); + session_teardown(s, "limit-exceeded"); + return (0); +} + /* RFC 9051 SS9: a tag may not contain "+" */ static int tag_is_valid(const char *tag) @@ -933,7 +1071,6 @@ session_dispatch_client(int fd, short event, void *arg size_t consumed, linelen; int alive; - /* a refused non-synchronizing literal's octets are in flight */ if (s->literal_discard > 0) { uint64_t take; @@ -947,12 +1084,7 @@ session_dispatch_client(int fd, short event, void *arg } if (s->literal_discard > 0) break; /* need more data */ - if (s->inbuflen >= 2 && s->inbuf[0] == '\r' && - s->inbuf[1] == '\n') { - memmove(s->inbuf, s->inbuf + 2, - s->inbuflen - 2); - s->inbuflen -= 2; - } + s->literal_skipline = 1; continue; } @@ -1001,6 +1133,30 @@ session_dispatch_client(int fd, short event, void *arg continue; } + if (s->cmd_gather && s->cmd_octets > 0) { + size_t take; + + take = s->inbuflen < s->cmd_octets ? + s->inbuflen : s->cmd_octets; + if (take == 0) + break; /* need more data */ + if (memchr(s->inbuf, '\0', take) != NULL) { + /* RFC 9051 SS9: CHAR8 excludes NUL */ + session_cmd_reply(s, "BAD", "NUL in a literal"); + s->literal_discard = s->cmd_octets; + s->cmd_gather = 0; + s->cmd_queued = 0; + s->cmd_octets = 0; + continue; + } + memcpy(s->cmdbuf + s->cmdlen, s->inbuf, take); + s->cmdlen += take; + s->cmd_octets -= take; + memmove(s->inbuf, s->inbuf + take, s->inbuflen - take); + s->inbuflen -= take; + continue; + } + crlf = memmem(s->inbuf, s->inbuflen, "\r\n", 2); if (crlf == NULL) break; @@ -1027,19 +1183,43 @@ session_dispatch_client(int fd, short event, void *arg nonsync_len = 0; if (line_nonsync_literal(s->inbuf, &nonsync_len) && nonsync_len > IMAP_NONSYNC_LITERAL_MAX) { - static const char bad[] = "* BAD non-synchronizing " - "literal exceeds 4096 octets, closing " - "connection\r\n"; + /* RFC 7888 SS4, SS5: BYE, with RFC 4469's TOOBIG */ + static const char bye[] = "* BYE [TOOBIG] " + "non-synchronizing literal exceeds 4096 octets\r\n"; log_warnx("session %u: oversized non-synchronizing " "literal (%llu), closing", s->id, (unsigned long long)nonsync_len); - session_write(s, bad, sizeof(bad) - 1); + session_write(s, bye, sizeof(bye) - 1); session_teardown(s, "limit-exceeded"); return; } - if (s->auth_cont) { + if (s->literal_skipline) { + s->literal_skipline = 0; + alive = 1; + } else if (s->cmd_gather) { + s->cmd_gather = 0; + alive = 1; + if (linelen >= sizeof(s->cmdbuf) - s->cmdlen) { + session_cmd_reply(s, "NO", cmd_too_long); + s->cmd_queued = 0; + } else { + memcpy(s->cmdbuf + s->cmdlen, s->inbuf, + linelen + 1); + if (!s->cmd_queued) + alive = session_run_cmd(s, s->cmdbuf); + else if (nonsync_len > 0) + (void)session_gather(s, nonsync_len); + else { + s->cmd_queued = 0; + alive = session_queue(s, s->cmdbuf); + } + } + if (alive && nonsync_len == 0 && !s->cmd_gather && + s->cmd_queue_n > 0 && !session_is_busy(s)) + alive = session_dequeue_next(s); + } else if (s->auth_cont) { /* the line is the user's base64 password, see below */ s->scrub_inbuf = 1; alive = session_handle_auth_continuation(s, s->inbuf); @@ -1047,31 +1227,23 @@ session_dispatch_client(int fd, short event, void *arg alive = session_handle_idle_continuation(s, s->inbuf); } else if (session_is_busy(s)) { /* RFC 9051 SS5.5: queue while one is in flight */ + alive = 1; if (nonsync_len > 0) { - session_reply(s, "*", "BAD", - "non-synchronizing literal not accepted on " - "a pipelined command, retry with a " - "synchronizing literal"); - alive = 1; - } else if (!session_enqueue_cmd(s, s->inbuf)) { - static const char bad[] = "* BAD too many " - "pipelined commands, closing connection" - "\r\n"; - - log_warnx("session %u: pipelined command " - "queue full, closing", s->id); - session_write(s, bad, sizeof(bad) - 1); - session_teardown(s, "limit-exceeded"); - return; + (void)strlcpy(s->cmdbuf, s->inbuf, + sizeof(s->cmdbuf)); + s->cmdlen = strlen(s->cmdbuf); + s->cmd_queued = 1; + (void)session_gather(s, nonsync_len); } else - alive = 1; + alive = session_queue(s, s->inbuf); } else { - alive = session_handle_line(s, s->inbuf); + alive = session_run_cmd(s, s->inbuf); } if (alive == 0) return; /* s was torn down (LOGOUT), do not touch */ - if (nonsync_len > 0 && !s->literal_pending) + if (nonsync_len > 0 && !s->literal_pending && + !s->cmd_gather) s->literal_discard = nonsync_len; /* cmd_starttls() zeroes inbuflen */ @@ -1371,7 +1543,8 @@ session_enqueue_cmd(struct session *s, const char *lin int session_dequeue_next(struct session *s) { - while (s->cmd_queue_n > 0 && !session_is_busy(s)) { + while (s->cmd_queue_n > 0 && !session_is_busy(s) && + !s->cmd_gather) { char *line = s->cmd_queue[0]; uint32_t i; int alive; @@ -1380,7 +1553,7 @@ session_dequeue_next(struct session *s) s->cmd_queue[i - 1] = s->cmd_queue[i]; s->cmd_queue_n--; - alive = session_handle_line(s, line); + alive = session_run_cmd(s, line); free(line); if (!alive) return (0); blob - 0ba65946e76d2bf2869f7b2e0ec1b71892492b88 blob + f33cba5aa3d76a8ba44c5945fc1a086bc3970e2d --- src/listener.h +++ src/listener.h @@ -140,8 +140,16 @@ struct session { int literal_pending; uint64_t literal_len; /* announced size, "{n}" */ uint64_t literal_remaining; - /* a refused non-synchronizing literal's octets, swallowed */ + /* RFC 7888 SS3: a refused literal's octets, then its command's line */ uint64_t literal_discard; + int literal_skipline; + /* RFC 9051 SS4.3: a command gathered with its literals */ + char cmdbuf[SESSION_INBUF_MAX]; + size_t cmdlen; + size_t cmd_octets; + int cmd_gather; + int cmd_queued; /* for the pipeline */ + int literal_wanted; char append_mailbox[MBOX_NAME_MAX]; enum session_state append_prev_state; @@ -230,6 +238,7 @@ extern struct tls *listener_tls_ctx; extern uint32_t listener_idle_poll_secs; extern uint32_t listener_login_grace_secs; extern uint64_t listener_append_max; +extern const char literal_wanted[]; void listener_dispatch_auth(int, short, void *); void listener_dispatch_parent(int, short, void *); @@ -344,6 +353,8 @@ int listener_reject_bad_utf8(struct session *, const int list_pattern_match(const char *, const char *, int); int parse_mailbox_name(char **, char *, size_t, const char **); int parse_list_pattern(char **, char *, size_t, const char **); +int read_literal(char **, char **, size_t *, const char **); +int session_arg_error(struct session *, const char *, const char *); int quote_mailbox(char *, size_t, const char *); void session_finish_mbox_op(struct session *, const struct imsg_mbox_result *); blob - 1536d4ddcfe4de80dc7e9f9403c651fae9aa611f blob + 7e8374d788e2868027027f32d27f31f14949c84f --- src/mailbox_cmd.c +++ src/mailbox_cmd.c @@ -271,10 +271,8 @@ select_or_examine(struct session *s, const char *tag, } p = args; - if (parse_mailbox_name(&p, mailbox, sizeof(mailbox), &errmsg) == -1) { - session_reply(s, tag, "BAD", errmsg); - return (1); - } + if (parse_mailbox_name(&p, mailbox, sizeof(mailbox), &errmsg) == -1) + return (session_arg_error(s, tag, errmsg)); while (*p == ' ') p++; params = (*p != '\0') ? p : NULL; @@ -395,10 +393,8 @@ cmd_create(struct session *s, const char *tag, char *a } p = args; - if (parse_mailbox_name(&p, mailbox, sizeof(mailbox), &errmsg) == -1) { - session_reply(s, tag, "BAD", errmsg); - return (1); - } + if (parse_mailbox_name(&p, mailbox, sizeof(mailbox), &errmsg) == -1) + return (session_arg_error(s, tag, errmsg)); if (mailbox_name_is_inbox(mailbox)) { session_reply(s, tag, "NO", "[CANNOT] cannot create INBOX"); @@ -467,10 +463,8 @@ cmd_delete(struct session *s, const char *tag, char *a } p = args; - if (parse_mailbox_name(&p, mailbox, sizeof(mailbox), &errmsg) == -1) { - session_reply(s, tag, "BAD", errmsg); - return (1); - } + if (parse_mailbox_name(&p, mailbox, sizeof(mailbox), &errmsg) == -1) + return (session_arg_error(s, tag, errmsg)); if (mailbox_name_is_inbox(mailbox)) { session_reply(s, tag, "NO", "[CANNOT] cannot delete INBOX"); @@ -538,14 +532,10 @@ cmd_rename(struct session *s, const char *tag, char *a } p = args; - if (parse_mailbox_name(&p, oldname, sizeof(oldname), &errmsg) == -1) { - session_reply(s, tag, "BAD", errmsg); - return (1); - } - if (parse_mailbox_name(&p, newname, sizeof(newname), &errmsg) == -1) { - session_reply(s, tag, "BAD", errmsg); - return (1); - } + if (parse_mailbox_name(&p, oldname, sizeof(oldname), &errmsg) == -1) + return (session_arg_error(s, tag, errmsg)); + if (parse_mailbox_name(&p, newname, sizeof(newname), &errmsg) == -1) + return (session_arg_error(s, tag, errmsg)); if (mailbox_name_is_inbox(oldname)) { /* RFC 9051 SS6.3.6 allows refusing this */ @@ -633,10 +623,8 @@ subscribe_dispatch(struct session *s, const char *tag, } p = args; - if (parse_mailbox_name(&p, mailbox, sizeof(mailbox), &errmsg) == -1) { - session_reply(s, tag, "BAD", errmsg); - return (1); - } + if (parse_mailbox_name(&p, mailbox, sizeof(mailbox), &errmsg) == -1) + return (session_arg_error(s, tag, errmsg)); if (listener_reject_bad_utf8(s, tag, mailbox)) return (1); @@ -812,9 +800,19 @@ parse_mailbox_arg(char **pp, char *out, size_t outsize } if (*p == '{') { - *errmsg = "literals are not supported in a mailbox name; " - "send it as a quoted string"; - return (-1); + char *lit; + size_t len; + + if (read_literal(&p, &lit, &len, errmsg) == -1) + return (-1); + if (len >= outsize) { + *errmsg = "mailbox name too long"; + return (-1); + } + memcpy(out, lit, len); + out[len] = '\0'; + *pp = p; + return (0); } if (*p == '"') { @@ -939,10 +937,8 @@ list_dispatch(struct session *s, const char *tag, char } if (parse_mailbox_name(&p, reference, sizeof(reference), &errmsg) == - -1) { - session_reply(s, tag, "BAD", errmsg); - return (1); - } + -1) + return (session_arg_error(s, tag, errmsg)); while (*p == ' ') p++; @@ -956,10 +952,8 @@ list_dispatch(struct session *s, const char *tag, char return (1); } - if (parse_list_pattern(&p, pattern, sizeof(pattern), &errmsg) == -1) { - session_reply(s, tag, "BAD", errmsg); - return (1); - } + if (parse_list_pattern(&p, pattern, sizeof(pattern), &errmsg) == -1) + return (session_arg_error(s, tag, errmsg)); while (*p == ' ') p++; @@ -1077,10 +1071,8 @@ cmd_status(struct session *s, const char *tag, char *a } p = args; - if (parse_mailbox_name(&p, mailbox, sizeof(mailbox), &errmsg) == -1) { - session_reply(s, tag, "BAD", errmsg); - return (1); - } + if (parse_mailbox_name(&p, mailbox, sizeof(mailbox), &errmsg) == -1) + return (session_arg_error(s, tag, errmsg)); while (*p == ' ') p++; blob - 03da0e60a5cf8a9436a9f42d69541fb632f44545 blob + 458f5c3e1ebc8d703f72a9e02f98902b7b6a33c1 --- src/mbox_copy.c +++ src/mbox_copy.c @@ -686,6 +686,7 @@ move_same_mailbox(struct imsg_mbox_copy *req, const st for (i = 0; i < nmoved; i++) { struct imsg_mbox_expunged exp; + pcache_drop_uid(ss->selected_mailbox, moved[i].old_uid); memset(&exp, 0, sizeof(exp)); exp.seqno = moved[i].old_seqno; exp.uid = moved[i].old_uid; @@ -756,6 +757,7 @@ move_cross_mailbox(struct imsg_mbox_copy *req, const s continue; } any_removed = 1; + pcache_drop_uid(ss->selected_mailbox, staged[i].src_uid); if (locate_message_file(&ss->cur_snap, ss->mailbox_dir_fd, rec.basename, &size, suffix, sizeof(suffix)) == 0) { blob - a0ae81162de65bd2b4a9878669c2f7cdd5292754 blob + 0b694205c1d879f3ca71a8c24865436b4b7da1e7 --- src/mbox_fetch.c +++ src/mbox_fetch.c @@ -38,6 +38,8 @@ static void fetch_walk_step(struct store_session *); static void fetch_walk_finish(struct store_session *, int); +static const char *fetch_cached(struct store_session *, + const struct index_rec *, int, int *, uint32_t *); static void fetch_send_part(struct store_session *ss, int imsg_type, const char *what, @@ -98,6 +100,7 @@ handle_mbox_fetch(struct imsg_mbox_fetch *req, const s return (0); } index_lock_release(&il); + pcache_check_mailbox(ss->selected_mailbox, idx->uidvalidity); /* RFC 9051 SS6.4.9 */ fw->nresolved = seqset_resolve(ranges, nranges, req->by_uid ? @@ -118,10 +121,15 @@ handle_mbox_fetch(struct imsg_mbox_fetch *req, const s } static int -fetch_needs_parser(const struct imsg_mbox_fetch *req) +fetch_needs_parser(const struct imsg_mbox_fetch *req, const char *mailbox, + uint32_t uidvalidity, const struct index_rec *rec) { - if (req->attrs & (MBOX_FETCH_ENVELOPE | MBOX_FETCH_BODYSTRUCTURE)) + if ((req->attrs & MBOX_FETCH_ENVELOPE) && !pcache_has(mailbox, + uidvalidity, rec->uid, rec->basename, PCACHE_ENVELOPE)) return (1); + if ((req->attrs & MBOX_FETCH_BODYSTRUCTURE) && !pcache_has(mailbox, + uidvalidity, rec->uid, rec->basename, PCACHE_BODYSTRUCTURE)) + return (1); if ((req->attrs & MBOX_FETCH_HEADER_FIELDS) && !(req->attrs & MBOX_FETCH_BODY_HEADER)) return (1); @@ -129,6 +137,26 @@ fetch_needs_parser(const struct imsg_mbox_fetch *req) !(req->attrs & (MBOX_FETCH_BODY_WHOLE | MBOX_FETCH_BODY_TEXT))); } +/* a hit still needs the file on disk, as open_message_file() does */ +static const char * +fetch_cached(struct store_session *ss, const struct index_rec *rec, + int item, int *have_file, uint32_t *lenp) +{ + const char *text; + off_t size; + char suffix[64]; + + if ((text = pcache_get(ss->selected_mailbox, ss->fetch.idx.uidvalidity, + rec->uid, rec->basename, item, lenp)) == NULL) + return (NULL); + if (*have_file == 0 && locate_message_file(&ss->cur_snap, + ss->mailbox_dir_fd, rec->basename, &size, suffix, + sizeof(suffix)) == -1) + return (NULL); + *have_file = 1; + return (text); +} + /* returns, still active, once FETCH_BATCH_MAX is queued */ static void fetch_walk_step(struct store_session *ss) @@ -163,7 +191,8 @@ fetch_walk_step(struct store_session *ss) if (req->has_changedsince && rec.modseq <= req->changedsince) continue; - if (!fw->no_parser && fetch_needs_parser(req)) { + if (!fw->no_parser && fetch_needs_parser(req, + ss->selected_mailbox, idx->uidvalidity, &rec)) { int ready = parser_ready(); if (ready == 0) { @@ -333,19 +362,28 @@ fetch_walk_step(struct store_session *ss) if (req->attrs & MBOX_FETCH_ENVELOPE) { struct imsg_mbox_fetch_envelope envmeta; struct imsg_parser_rep prep; + const char *env; char *envbuf = NULL; int mfd; memset(&envmeta, 0, sizeof(envmeta)); envmeta.seqno = i; envmeta.uid = rec.uid; - if ((mfd = open_message_file(&ss->cur_snap, + if ((env = fetch_cached(ss, &rec, PCACHE_ENVELOPE, + &have_file, &envmeta.envlen)) != NULL) + envmeta.found = 1; + else if ((mfd = open_message_file(&ss->cur_snap, ss->mailbox_dir_fd, rec.basename)) != -1) { if (parser_request(IMSG_PARSER_ENVELOPE, "ENVELOPE", mfd, rec.basename, NULL, NULL, 0, ENVELOPE_MAX, &prep, &envbuf) == 0) { envmeta.found = 1; envmeta.envlen = prep.len; + env = envbuf; + pcache_put(ss->selected_mailbox, + idx->uidvalidity, rec.uid, + rec.basename, PCACHE_ENVELOPE, + envbuf, prep.len); } close(mfd); } @@ -353,20 +391,24 @@ fetch_walk_step(struct store_session *ss) fetch_send_part(ss, IMSG_MBOX_FETCH_ENVELOPE, "IMSG_MBOX_FETCH_ENVELOPE", &envmeta, sizeof(envmeta), - envmeta.found, envbuf, envmeta.envlen); + envmeta.found, env, envmeta.envlen); free(envbuf); } if (req->attrs & MBOX_FETCH_BODYSTRUCTURE) { struct imsg_mbox_fetch_bodystructure bsmeta; struct imsg_parser_rep prep; + const char *bs; char *bsbuf = NULL; int mfd; memset(&bsmeta, 0, sizeof(bsmeta)); bsmeta.seqno = i; bsmeta.uid = rec.uid; - if ((mfd = open_message_file(&ss->cur_snap, + if ((bs = fetch_cached(ss, &rec, PCACHE_BODYSTRUCTURE, + &have_file, &bsmeta.bslen)) != NULL) + bsmeta.found = 1; + else if ((mfd = open_message_file(&ss->cur_snap, ss->mailbox_dir_fd, rec.basename)) != -1) { if (parser_request(IMSG_PARSER_BODYSTRUCTURE, "BODYSTRUCTURE", mfd, rec.basename, NULL, @@ -374,13 +416,18 @@ fetch_walk_step(struct store_session *ss) &bsbuf) == 0) { bsmeta.found = 1; bsmeta.bslen = prep.len; + bs = bsbuf; + pcache_put(ss->selected_mailbox, + idx->uidvalidity, rec.uid, + rec.basename, PCACHE_BODYSTRUCTURE, + bsbuf, prep.len); } close(mfd); } fetch_send_part(ss, IMSG_MBOX_FETCH_BODYSTRUCTURE, "IMSG_MBOX_FETCH_BODYSTRUCTURE", &bsmeta, - sizeof(bsmeta), bsmeta.found, bsbuf, bsmeta.bslen); + sizeof(bsmeta), bsmeta.found, bs, bsmeta.bslen); free(bsbuf); } blob - 1dc72a9e2541dae6be222c72dbf7351f2615bd57 blob + 63953515d197c7229551665b077361c892fc219a --- src/mbox_manage.c +++ src/mbox_manage.c @@ -359,6 +359,8 @@ handle_mbox_delete(struct imsg_mbox_delete *req, struc result.error = MBOX_OP_OK; send: + if (result.error == MBOX_OP_OK) + pcache_drop_mailbox(req->mailbox); /* deleting the selection clears the store's gate */ if (result.error == MBOX_OP_OK && strcmp(ss->selected_mailbox, req->mailbox) == 0) @@ -421,6 +423,8 @@ handle_mbox_rename(struct imsg_mbox_rename *req, struc result.error = MBOX_OP_OK; send: + if (result.error == MBOX_OP_OK) + pcache_drop_mailbox(req->oldname); if (result.error == MBOX_OP_OK && strcmp(ss->selected_mailbox, req->oldname) == 0) { if (strlcpy(ss->selected_mailbox, req->newname, blob - 77c3a0fc50453fb1d0601bbdf426d29413f81417 blob + 3f440822a5f3f8e4528335336b76cdb09c52bb88 --- src/mbox_search.c +++ src/mbox_search.c @@ -216,7 +216,8 @@ search_op_content(int op) return (op == SEARCH_OP_SUBJECT || op == SEARCH_OP_HEADER || op == SEARCH_OP_SENTBEFORE || op == SEARCH_OP_SENTON || op == SEARCH_OP_SENTSINCE || op == SEARCH_OP_FROM || - op == SEARCH_OP_TO || op == SEARCH_OP_CC || op == SEARCH_OP_BCC); + op == SEARCH_OP_TO || op == SEARCH_OP_CC || op == SEARCH_OP_BCC || + op == SEARCH_OP_BODY || op == SEARCH_OP_TEXT); } static int blob - 281b003d5dc7a3ec06075dc3f18e345f7cf3e01a blob + ffb4cf902c1fc490f7cdf4d038efe4efab3c46f3 --- src/mbox_store.c +++ src/mbox_store.c @@ -460,6 +460,7 @@ handle_mbox_expunge(struct imsg_mbox_expunge *req, if (index_parse_line(gone[k].line, &rec) == -1) continue; + pcache_drop_uid(ss->selected_mailbox, rec.uid); if (snprintf(path, sizeof(path), "%s/%s%s", gone[k].suffix[0] == '\0' ? "new" : "cur", rec.basename, gone[k].suffix) >= (int)sizeof(path)) blob - fa845ff835f6518c9e1430ea20524343597c0f2b blob + b27d4cefa382bca517e4f751af19673567f77a9d --- src/mime.c +++ src/mime.c @@ -433,6 +433,7 @@ header_field_name_matches(const char *name, size_t nam /* one field plus its RFC 5322 SS2.2.3 obs-fold lines */ struct hdr_field { size_t start; + size_t name_end; size_t colon; size_t line_end; size_t end; @@ -459,7 +460,7 @@ hdr_next_field(const char *hdr, size_t hdrlen, size_t *off = i + 1; if (f->line_end == f->start) { - f->colon = f->line_end; + f->colon = f->name_end = f->line_end; f->end = *off; return (0); } @@ -468,6 +469,11 @@ hdr_next_field(const char *hdr, size_t hdrlen, size_t if (hdr[f->colon] == ':') break; } + /* RFC 5322 SS4.5: white space may precede the colon */ + f->name_end = f->colon; + while (f->colon < f->line_end && f->name_end > f->start && + (hdr[f->name_end - 1] == ' ' || hdr[f->name_end - 1] == '\t')) + f->name_end--; while (*off < hdrlen && (hdr[*off] == ' ' || hdr[*off] == '\t')) { j = *off; @@ -511,7 +517,7 @@ filter_header_fields(const char *hdr, size_t hdrlen, c } matched = header_field_name_matches(hdr + f.start, - f.colon - f.start, fields_spec); + f.name_end - f.start, fields_spec); include = want_not ? !matched : matched; if (include) { @@ -578,7 +584,7 @@ extract_header_field(const char *hdr, size_t hdrlen, c if (hdr_next_field(hdr, hdrlen, &off, &f) != 1) return (-1); - if (f.colon - f.start != namelen || + if (f.name_end - f.start != namelen || strncasecmp(hdr + f.start, name, namelen) != 0) continue; @@ -633,7 +639,7 @@ header_next_field(const char *hdr, size_t hdrlen, size (hdr[vstart] == ' ' || hdr[vstart] == '\t')) vstart++; *name = hdr + f.start; - *namelen = f.colon - f.start; + *namelen = f.name_end - f.start; *val = hdr + vstart; *vallen = f.end - vstart; return (1); @@ -695,9 +701,9 @@ mime_read_token_or_qstring(const char *s, size_t len, (*pos)++; c = s[*pos]; } - /* no CR, LF or NUL inside a quoted-string either */ - if (c == '\0' || c == '\r' || c == '\n') - return (-1); + /* RFC 5322 SS4: keep CR and LF; a NUL would end out */ + if (c == '\0') + c = ' '; if (outlen + 1 >= outsize) return (-1); out[outlen++] = c; blob - 1ae52cd9d5eeda56251857220f67e5fde5d6bb59 blob + 7afa4a6273c203c4638a8f55ff478bdda2a2aeb0 --- src/search_cmd.c +++ src/search_cmd.c @@ -84,7 +84,6 @@ struct search_parse_ctx { int uses_modseq; /* octets of every string operand, counted past the cap too */ size_t operand_len; - int uses_content; char pool[SEARCH_OPERANDS_MAX]; }; @@ -257,11 +256,14 @@ read_search_astring(char **pp, struct search_parse_ctx while (*p == ' ') p++; if (*p == '{') { - *errmsg = "literals are not supported in SEARCH, send a " - "quoted string"; - return (-1); - } - if (*p == '"') { + char *lit; + size_t len, i; + + if (read_literal(&p, &lit, &len, errmsg) == -1) + return (-1); + for (i = 0; i < len; i++) + search_pool_add(ctx, lit[i]); + } else if (*p == '"') { for (p++; *p != '"'; p++) { if (*p == '\0') { *errmsg = "unterminated quoted string"; @@ -617,10 +619,12 @@ parse_search_key_inner(char **pp, struct search_parse_ int op; } string_keys[] = { { "BCC", SEARCH_OP_BCC }, + { "BODY", SEARCH_OP_BODY }, { "CC", SEARCH_OP_CC }, { "FROM", SEARCH_OP_FROM }, { "HEADER", SEARCH_OP_HEADER }, { "SUBJECT", SEARCH_OP_SUBJECT }, + { "TEXT", SEARCH_OP_TEXT }, { "TO", SEARCH_OP_TO }, }; struct search_node node; @@ -648,17 +652,6 @@ parse_search_key_inner(char **pp, struct search_parse_ } } - /* BODY and TEXT: parsed, but not yet answered */ - if (strcasecmp(word, "BODY") == 0 || strcasecmp(word, "TEXT") == 0) { - uint32_t off, len; - - if (read_search_astring(&p, ctx, &off, &len, errmsg) == -1) - return (-1); - ctx->uses_content = 1; - *pp = p; - return (0); - } - *errmsg = "unknown search key"; return (-1); } @@ -730,10 +723,6 @@ search_program_parse(char *args, struct search_node *n if (rc == 0 && ctx.operand_len > SEARCH_OPERANDS_MAX) { errmsg = "[LIMIT] search strings exceed 4096 octets in all"; rc = -2; - } else if (rc == 0 && ctx.uses_content) { - errmsg = "search keys that require message content/header " - "access are not supported in this pass"; - rc = -2; } else if (rc == 0 && ctx.n == 0) { errmsg = "SEARCH requires search criteria"; rc = -1; @@ -814,6 +803,43 @@ parse_search_return_opts(char **pp, uint32_t *opts_out } } +/* RFC 9051 SS9 charset: atom / quoted */ +static int +read_search_charset(char **pp, char *out, size_t outsize) +{ + char *p = *pp; + size_t o = 0; + + if (*p == '"') { + for (p++; *p != '"'; p++) { + if (*p == '\0' || + (*p == '\\' && *++p != '"' && *p != '\\')) + return (-1); + if (o + 1 >= outsize) + return (-1); + out[o++] = *p; + } + p++; + } else { + /* ATOM-CHAR: any CHAR but atom-specials */ + for (; *p != '\0' && *p != ' '; p++) { + if ((unsigned char)*p <= 0x1f || + (unsigned char)*p >= 0x7f || + strchr("(){%*\"\\]", *p) != NULL || + o + 1 >= outsize) + return (-1); + out[o++] = *p; + } + if (o == 0) + return (-1); + } + if (*p != ' ' && *p != '\0') + return (-1); + out[o] = '\0'; + *pp = p; + return (0); +} + static void search_dispatch_finish(struct session *, const struct search_parse_result *, struct search_node *); @@ -869,23 +895,15 @@ search_dispatch(struct session *s, const char *tag, ch if (strncasecmp(p, "CHARSET", 7) == 0 && (p[7] == ' ' || p[7] == '\0')) { - const char *start; char charset[64]; - size_t len; p += 7; while (*p == ' ') p++; - start = p; - while (*p != '\0' && *p != ' ') - p++; - len = (size_t)(p - start); - if (len == 0 || len >= sizeof(charset)) { + if (read_search_charset(&p, charset, sizeof(charset)) == -1) { session_reply(s, tag, "BAD", "malformed CHARSET"); return (1); } - memcpy(charset, start, len); - charset[len] = '\0'; if (strcasecmp(charset, "US-ASCII") != 0 && strcasecmp(charset, "UTF-8") != 0) { @@ -912,19 +930,6 @@ search_dispatch(struct session *s, const char *tag, ch return (1); } - free(s->search_matches); - s->search_matches = NULL; - s->search_nmatches = 0; - s->search_matches_cap = 0; - s->search_alloc_failed = 0; - s->search_return_opts = return_opts; - s->cmd_by_uid = by_uid; - - if (strlcpy(s->pending_tag, tag, sizeof(s->pending_tag)) >= - sizeof(s->pending_tag)) { - session_reply(s, tag, "NO", "[SERVERBUG] internal error"); - return (1); - } { struct search_parse_result res; struct search_node nodes[SEARCH_PROGRAM_MAX_NODES]; @@ -933,6 +938,23 @@ search_dispatch(struct session *s, const char *tag, ch res.rc = search_program_parse(p, nodes, &res.nnodes, &res.uses_modseq, res.pool, &res.poollen, res.errmsg, sizeof(res.errmsg)); + if (res.rc == -1 && strcmp(res.errmsg, literal_wanted) == 0) + return (session_arg_error(s, tag, literal_wanted)); + + free(s->search_matches); + s->search_matches = NULL; + s->search_nmatches = 0; + s->search_matches_cap = 0; + s->search_alloc_failed = 0; + s->search_return_opts = return_opts; + s->cmd_by_uid = by_uid; + + if (strlcpy(s->pending_tag, tag, sizeof(s->pending_tag)) >= + sizeof(s->pending_tag)) { + session_reply(s, tag, "NO", + "[SERVERBUG] internal error"); + return (1); + } search_dispatch_finish(s, &res, res.nnodes > 0 ? nodes : NULL); } blob - f514c63455866a48bef14a150964c849578eced7 blob + 83080cc8b4803a2e5695bc813d3e09df2b69d5dd --- src/search_match.c +++ src/search_match.c @@ -23,6 +23,7 @@ #include #include #include +#include #include #include #include @@ -34,6 +35,29 @@ #define SEARCH_READ_BLOCK 65536 +#define PART_SKIP 0 +#define PART_TEXT 1 +#define PART_MULTI 2 +#define PART_MESSAGE 3 + +#define ENC_IDENTITY 0 +#define ENC_BASE64 1 +#define ENC_QP 2 + +struct body_key { + int op; + char *fold; + size_t *fail; + size_t len; + char *hit; +}; + +struct body_walk { + struct body_key *keys; + uint32_t nkeys; + uint32_t left; +}; + static const char *const month_names[12] = { "Jan", "Feb", "Mar", "Apr", "May", "Jun", "Jul", "Aug", "Sep", "Oct", "Nov", "Dec", @@ -48,11 +72,29 @@ static size_t encoded_word_len(const char *, size_t); static char *decode_value(const char *, size_t, size_t *); static int date_fws(int); static int date_field_day(const char *, size_t, int64_t *); -static int read_header(int, const char *, char **, size_t *); -static char *unescape_name(const char *, size_t, size_t *); +static int read_header(int, const char *, int, char **, size_t *); static int match_addresses(const char *, size_t, const char *, size_t); static int match_leaf(const char *, size_t, const struct imsg_parser_leaf *, const char *); +static void kmp_prepare(const char *, size_t, char *, size_t *); +static int kmp_contains(const char *, size_t, const char *, + const size_t *, size_t); +static size_t decode_b64(const char *, size_t, char *); +static size_t decode_qp(const char *, size_t, char *); +static void scan_text(struct body_walk *, const char *, size_t, int); +static int scan_header(struct body_walk *, const char *, size_t); +static int scan_body(struct body_walk *, const char *, size_t, int); +static int part_type(const char *, size_t, int, char *, size_t, + int *, int *); +static int next_delimiter(const char *, size_t, const char *, + size_t, size_t *, size_t *, int *); +static int walk_multipart(struct body_walk *, int, const char *, + size_t, const char *, int); +static int walk_entity(struct body_walk *, int, const char *, size_t, + const char *, size_t, int); +static int match_body(const char *, size_t, size_t, + const struct imsg_parser_leaf *, uint32_t, const char *, + char *); /* RFC 9051 SS6.4.4 folds the ASCII range only, byte by byte */ static int @@ -299,8 +341,10 @@ date_field_day(const char *v, size_t vlen, int64_t *ou return (0); } +/* the header, or with whole set the message, up to attachment max */ static int -read_header(int fd, const char *basename, char **buf_out, size_t *len_out) +read_header(int fd, const char *basename, int whole, char **buf_out, + size_t *len_out) { char *buf = NULL, *nbuf; size_t len = 0, size = 0, hdrend, limit; @@ -312,8 +356,9 @@ read_header(int fd, const char *basename, char **buf_o if (len == size) { if (size == limit) break; - size = (limit - size > SEARCH_READ_BLOCK) ? - size + SEARCH_READ_BLOCK : limit; + size = size == 0 ? SEARCH_READ_BLOCK : size * 2; + if (size > limit) + size = limit; if ((nbuf = realloc(buf, size)) == NULL) { log_warn("session %u: realloc SEARCH header", session_id); @@ -333,14 +378,15 @@ read_header(int fd, const char *basename, char **buf_o if (n == 0) break; len += (size_t)n; - if (find_header_body_split(buf, len, &hdrend) == 0) { + if (!whole && find_header_body_split(buf, len, &hdrend) == 0) { len = hdrend; break; } } if (len > bodystructure_read_max) { - log_warnx("session %u: message %s has no header end within " - "%u bytes, SEARCH cannot check it", session_id, basename, + log_warnx("session %u: message %s %s %u bytes, SEARCH " + "cannot check it", session_id, basename, whole ? + "is over" : "has no header end within", bodystructure_read_max); free(buf); return (-1); @@ -350,34 +396,15 @@ read_header(int fd, const char *basename, char **buf_o return (0); } -/* An RFC 5322 SS3.2.4 quoted string's content with its quoted-pairs undone */ -static char * -unescape_name(const char *v, size_t vlen, size_t *outlen) -{ - char *out; - size_t i, n = 0; - - if ((out = malloc(vlen + 1)) == NULL) - return (NULL); - for (i = 0; i < vlen; i++) { - if (v[i] == '\\' && i + 1 < vlen) - i++; - out[n++] = v[i]; - } - *outlen = n; - return (out); -} - -/* RFC 9051 SS6.4.4: addresses as ENVELOPE shows them */ +/* RFC 9051 SS6.4.4: addresses as ENVELOPE shows them, not group names */ static int match_addresses(const char *v, size_t vlen, const char *needle, size_t needlelen) { - const char *tok, *name, *mbox, *host; - char *unf, *spec, *raw, *dec; - size_t i, ulen = 0, pos = 0, toklen, namelen, mboxlen; - size_t hostlen, rawlen, declen; - int hit = 0; + struct address a; + char *unf, *tok, *dec; + size_t i, ulen = 0, pos = 0, toklen, declen; + int hit = 0, ingroup = 0, kind; /* RFC 5322 SS2.2.3: unfolding drops CR and LF */ if ((unf = malloc(vlen + 1)) == NULL) @@ -386,34 +413,24 @@ match_addresses(const char *v, size_t vlen, const char if (v[i] != '\r' && v[i] != '\n') unf[ulen++] = v[i]; } + ulen = address_uncomment(unf, ulen); - while (hit == 0 && address_list_next(unf, ulen, &pos, &tok, - &toklen) == 1) { - if (address_split(tok, toklen, &name, &namelen, &mbox, - &mboxlen, &host, &hostlen) == -1) + while (hit == 0 && (kind = address_list_next(unf, ulen, &pos, + &ingroup, &tok, &toklen)) != 0) { + if (kind != ADDR_MAILBOX || + address_split(tok, toklen, &a) == -1) continue; - if ((spec = malloc(mboxlen + 1 + hostlen)) == NULL) { - hit = -1; - break; - } - memcpy(spec, mbox, mboxlen); - spec[mboxlen] = '@'; - memcpy(spec + mboxlen + 1, host, hostlen); - hit = ci_contains(spec, mboxlen + 1 + hostlen, needle, - needlelen); - free(spec); - if (hit || name == NULL) + /* cooking only shortens, so mailbox@host fits where it was */ + a.mailbox[a.mailboxlen] = '@'; + memmove(a.mailbox + a.mailboxlen + 1, a.host, a.hostlen); + hit = ci_contains(a.mailbox, a.mailboxlen + 1 + a.hostlen, + needle, needlelen); + if (hit || a.name == NULL) continue; - if ((raw = unescape_name(name, namelen, &rawlen)) == NULL) { + if ((dec = decode_value(a.name, a.namelen, &declen)) == NULL) { hit = -1; break; } - dec = decode_value(raw, rawlen, &declen); - free(raw); - if (dec == NULL) { - hit = -1; - break; - } hit = ci_contains(dec, declen, needle, needlelen); free(dec); } @@ -500,25 +517,449 @@ match_leaf(const char *hdr, size_t hdrlen, return (hit); } +/* Knuth-Morris-Pratt, so a crafted body and needle cost linear time */ +static void +kmp_prepare(const char *needle, size_t len, char *fold, size_t *fail) +{ + size_t i, k = 0; + + for (i = 0; i < len; i++) + fold[i] = (char)ascii_lower((unsigned char)needle[i]); + if (len > 0) + fail[0] = 0; + for (i = 1; i < len; i++) { + while (k > 0 && fold[i] != fold[k]) + k = fail[k - 1]; + if (fold[i] == fold[k]) + k++; + fail[i] = k; + } +} + +static int +kmp_contains(const char *hay, size_t haylen, const char *fold, + const size_t *fail, size_t len) +{ + size_t i, k = 0; + int c; + + if (len == 0) + return (1); + for (i = 0; i < haylen; i++) { + c = ascii_lower((unsigned char)hay[i]); + while (k > 0 && c != (unsigned char)fold[k]) + k = fail[k - 1]; + if (c == (unsigned char)fold[k] && ++k == len) + return (1); + } + return (0); +} + +/* RFC 2045 SS6.8: characters outside the alphabet are ignored */ +static size_t +decode_b64(const char *in, size_t len, char *out) +{ + size_t i, n = 0; + unsigned int bits = 0; + int nbits = 0, v; + + for (i = 0; i < len && in[i] != '='; i++) { + if ((v = b64_value((unsigned char)in[i])) == -1) + continue; + bits = ((bits << 6) | (unsigned int)v) & 0xfff; + if ((nbits += 6) >= 8) { + nbits -= 8; + out[n++] = (char)((bits >> nbits) & 0xff); + } + } + return (n); +} + +/* RFC 2045 SS6.7; an "=" that is neither form is kept as it is */ +static size_t +decode_qp(const char *in, size_t len, char *out) +{ + size_t i, j, n = 0; + int hi, lo; + + for (i = 0; i < len; i++) { + if (in[i] != '=') { + out[n++] = in[i]; + continue; + } + for (j = i + 1; j < len && (in[j] == ' ' || in[j] == '\t'); + j++) + ; + if (j < len && in[j] == '\n') { + i = j; + continue; + } + if (j + 1 < len && in[j] == '\r' && in[j + 1] == '\n') { + i = j + 1; + continue; + } + if (i + 2 < len && + (hi = hex_value((unsigned char)in[i + 1])) != -1 && + (lo = hex_value((unsigned char)in[i + 2])) != -1) { + out[n++] = (char)(hi << 4 | lo); + i += 2; + continue; + } + out[n++] = in[i]; + } + return (n); +} + +static void +scan_text(struct body_walk *w, const char *t, size_t len, int header) +{ + struct body_key *k; + uint32_t i; + + for (i = 0; i < w->nkeys && w->left > 0; i++) { + k = &w->keys[i]; + if (*k->hit || (header && k->op != SEARCH_OP_TEXT)) + continue; + if (kmp_contains(t, len, k->fold, k->fail, k->len)) { + *k->hit = 1; + w->left--; + } + } +} + +/* RFC 9051 SS6.4.4: TEXT sees every header, MIME part headers too */ +static int +scan_header(struct body_walk *w, const char *hdr, size_t hdrlen) +{ + const char *name, *val; + char *dec, *line; + size_t namelen, vallen, dlen, off = 0; + uint32_t i; + + for (i = 0; i < w->nkeys; i++) { + if (*w->keys[i].hit == 0 && w->keys[i].op == SEARCH_OP_TEXT) + break; + } + if (i == w->nkeys) + return (0); + + while (w->left > 0 && header_next_field(hdr, hdrlen, &off, &name, + &namelen, &val, &vallen) == 1) { + if ((dec = decode_value(val, vallen, &dlen)) == NULL) + return (-1); + if ((line = malloc(namelen + 2 + dlen)) == NULL) { + free(dec); + return (-1); + } + memcpy(line, name, namelen); + memcpy(line + namelen, ": ", 2); + memcpy(line + namelen + 2, dec, dlen); + scan_text(w, line, namelen + 2 + dlen, 1); + free(line); + free(dec); + } + return (0); +} + +static int +scan_body(struct body_walk *w, const char *body, size_t len, int enc) +{ + char *out; + size_t n; + + if (enc == ENC_IDENTITY) { + scan_text(w, body, len, 0); + return (0); + } + /* neither decoding lengthens its input */ + if ((out = malloc(len + 1)) == NULL) + return (-1); + n = (enc == ENC_BASE64) ? decode_b64(body, len, out) : + decode_qp(body, len, out); + scan_text(w, out, n, 0); + free(out); + return (0); +} + +/* RFC 2045 SS5.2, SS6.4; RFC 2046 SS5.1.3, SS5.1.5 */ +static int +part_type(const char *hdr, size_t hdrlen, int digest, char *boundary, + size_t boundarysize, int *enc, int *subdigest) +{ + char *val = NULL, type[64], subtype[64], attr[64], value[256]; + size_t vallen = 0, pos = 0; + int kind, has_boundary = 0; + + *enc = ENC_IDENTITY; + *subdigest = 0; + if (extract_header_field(hdr, hdrlen, "Content-Transfer-Encoding", + &val, &vallen) == 0) { + while (vallen > 0 && (val[vallen - 1] == ' ' || + val[vallen - 1] == '\t')) + vallen--; + if (vallen == 6 && strncasecmp(val, "base64", 6) == 0) + *enc = ENC_BASE64; + else if (vallen == 16 && + strncasecmp(val, "quoted-printable", 16) == 0) + *enc = ENC_QP; + else if (!(vallen == 0 || + (vallen == 4 && (strncasecmp(val, "7bit", 4) == 0 || + strncasecmp(val, "8bit", 4) == 0)) || + (vallen == 6 && strncasecmp(val, "binary", 6) == 0))) { + /* SS6.4: an unknown one makes it octet-stream */ + free(val); + return (PART_SKIP); + } + free(val); + } + + /* SS5.2: none, or an invalid one, is text/plain */ + if (extract_header_field(hdr, hdrlen, "Content-Type", &val, + &vallen) == -1 || vallen == 0) { + free(val); + return (digest ? PART_MESSAGE : PART_TEXT); + } + if (mime_read_token_or_qstring(val, vallen, &pos, type, + sizeof(type)) == -1 || pos >= vallen || val[pos++] != '/' || + mime_read_token_or_qstring(val, vallen, &pos, subtype, + sizeof(subtype)) == -1) { + free(val); + return (PART_TEXT); + } + for (;;) { + while (pos < vallen && (val[pos] == ' ' || val[pos] == '\t')) + pos++; + if (pos >= vallen || val[pos++] != ';') + break; + while (pos < vallen && (val[pos] == ' ' || val[pos] == '\t')) + pos++; + if (mime_read_token_or_qstring(val, vallen, &pos, attr, + sizeof(attr)) == -1) + break; + while (pos < vallen && (val[pos] == ' ' || val[pos] == '\t')) + pos++; + if (pos >= vallen || val[pos++] != '=') + break; + while (pos < vallen && (val[pos] == ' ' || val[pos] == '\t')) + pos++; + if (mime_read_token_or_qstring(val, vallen, &pos, value, + sizeof(value)) == -1) + break; + /* RFC 2046 SS5.1.1: at most 70 characters */ + if (strcasecmp(attr, "boundary") == 0 && + strlcpy(boundary, value, boundarysize) < boundarysize) + has_boundary = 1; + } + free(val); + + if (strcasecmp(type, "multipart") == 0) { + /* RFC 2046 SS5.1.1 requires the boundary: invalid without */ + if (!has_boundary) + return (PART_TEXT); + *subdigest = strcasecmp(subtype, "digest") == 0; + kind = PART_MULTI; + } else if (strcasecmp(type, "message") == 0) + kind = (strcasecmp(subtype, "rfc822") == 0 || + strcasecmp(subtype, "global") == 0) ? PART_MESSAGE : + PART_TEXT; + else + kind = strcasecmp(type, "text") == 0 ? PART_TEXT : PART_SKIP; + /* RFC 2046 SS5.1.1, SS5.2.1: these are never encoded */ + if (kind != PART_TEXT) + *enc = ENC_IDENTITY; + return (kind); +} + +/* RFC 2046 SS5.1.1: the next "--boundary" line at or after *pos */ +static int +next_delimiter(const char *body, size_t len, const char *delim, + size_t dlen, size_t *pos, size_t *start, int *closing) +{ + const char *nl; + size_t i = *pos, after; + + while (i < len) { + if (len - i >= dlen && memcmp(body + i, delim, dlen) == 0) { + after = i + dlen; + *closing = 0; + if (len - after >= 2 && body[after] == '-' && + body[after + 1] == '-') { + *closing = 1; + after += 2; + } + while (after < len && + (body[after] == ' ' || body[after] == '\t')) + after++; + if (after == len || body[after] == '\n' || + (body[after] == '\r' && after + 1 < len && + body[after + 1] == '\n')) { + if (after < len) + after += body[after] == '\r' ? 2 : 1; + *start = i; + *pos = after; + return (1); + } + } + if ((nl = memchr(body + i, '\n', len - i)) == NULL) + break; + i = (size_t)(nl - body) + 1; + } + return (0); +} + +/* no part count limit: the message's size bounds it */ +static int +walk_multipart(struct body_walk *w, int depth, const char *body, + size_t len, const char *boundary, int digest) +{ + char delim[2 + 70 + 1]; + size_t dlen, pos = 0, dstart, pstart = 0, pend, phdr; + int closing, inpart = 0; + + dlen = (size_t)snprintf(delim, sizeof(delim), "--%s", boundary); + if (dlen >= sizeof(delim)) + return (0); + for (;;) { + int found = next_delimiter(body, len, delim, dlen, &pos, + &dstart, &closing); + + if (!found) + dstart = len; + if (inpart) { + /* the line end before a delimiter is the delimiter's */ + pend = dstart; + if (pend >= pstart + 2 && body[pend - 2] == '\r' && + body[pend - 1] == '\n') + pend -= 2; + else if (pend >= pstart + 1 && body[pend - 1] == '\n') + pend--; + if (pend == pstart) + phdr = 0; + else if (find_header_body_split(body + pstart, + pend - pstart, &phdr) == -1) + phdr = pend - pstart; + if (walk_entity(w, depth + 1, body + pstart, phdr, + body + pstart + phdr, pend - pstart - phdr, + digest) == -1) + return (-1); + } + /* a truncated multipart ends where the message does */ + if (!found || closing) + return (0); + pstart = pos; + inpart = 1; + } +} + +static int +walk_entity(struct body_walk *w, int depth, const char *hdr, size_t hdrlen, + const char *body, size_t bodylen, int digest) +{ + char boundary[70 + 1]; + size_t ih; + int enc, subdigest; + + if (w->left == 0) + return (0); + /* BODYSTRUCTURE's limit; deeper cannot be checked */ + if (depth > MIME_MAX_DEPTH) + return (-1); + if (scan_header(w, hdr, hdrlen) == -1) + return (-1); + switch (part_type(hdr, hdrlen, digest, boundary, sizeof(boundary), + &enc, &subdigest)) { + case PART_MULTI: + return (walk_multipart(w, depth, body, bodylen, boundary, + subdigest)); + case PART_MESSAGE: + if (bodylen == 0) + ih = 0; + else if (find_header_body_split(body, bodylen, &ih) == -1) + ih = bodylen; + return (walk_entity(w, depth + 1, body, ih, body + ih, + bodylen - ih, 0)); + case PART_TEXT: + return (scan_body(w, body, bodylen, enc)); + default: + return (0); + } +} + +/* RFC 9051 SS6.4.4 BODY and TEXT over TEXT and MESSAGE parts */ +static int +match_body(const char *msg, size_t hdrlen, size_t len, + const struct imsg_parser_leaf *leaves, uint32_t nleaves, + const char *pool, char *out) +{ + struct body_walk w; + struct body_key *k; + uint32_t i; + int rc = -1; + + memset(&w, 0, sizeof(w)); + if ((w.keys = calloc(nleaves, sizeof(*w.keys))) == NULL) + return (-1); + for (i = 0; i < nleaves; i++) { + if (leaves[i].op != SEARCH_OP_BODY && + leaves[i].op != SEARCH_OP_TEXT) + continue; + out[i] = 0; + /* "the empty string is a substring" */ + if (leaves[i].str_len == 0) { + out[i] = 1; + continue; + } + k = &w.keys[w.nkeys++]; + k->op = leaves[i].op; + k->len = leaves[i].str_len; + k->hit = &out[i]; + if ((k->fold = malloc(k->len)) == NULL || + (k->fail = calloc(k->len, sizeof(*k->fail))) == NULL) + goto done; + kmp_prepare(pool + leaves[i].str_off, k->len, k->fold, + k->fail); + w.left++; + } + rc = walk_entity(&w, 0, msg, hdrlen, msg + hdrlen, len - hdrlen, 0); +done: + for (i = 0; i < w.nkeys; i++) { + free(w.keys[i].fold); + free(w.keys[i].fail); + } + free(w.keys); + return (rc); +} + int search_match(int fd, const char *basename, const struct imsg_parser_leaf *leaves, uint32_t nleaves, const char *pool, char **buf_out, uint32_t *len_out) { char *hdr, *out; - size_t hdrlen; + size_t hdrlen, len; uint32_t i; - int m; + int m, whole = 0; *buf_out = NULL; *len_out = 0; - if (nleaves == 0 || read_header(fd, basename, &hdr, &hdrlen) == -1) + for (i = 0; i < nleaves; i++) { + if (leaves[i].op == SEARCH_OP_BODY || + leaves[i].op == SEARCH_OP_TEXT) + whole = 1; + } + if (nleaves == 0 || read_header(fd, basename, whole, &hdr, &len) == -1) return (-1); + /* a message with no blank line is all header */ + if (!whole || find_header_body_split(hdr, len, &hdrlen) == -1) + hdrlen = len; if ((out = malloc(nleaves)) == NULL) { free(hdr); return (-1); } for (i = 0; i < nleaves; i++) { + if (leaves[i].op == SEARCH_OP_BODY || + leaves[i].op == SEARCH_OP_TEXT) + continue; if ((m = match_leaf(hdr, hdrlen, &leaves[i], pool)) == -1) { free(out); free(hdr); @@ -526,6 +967,12 @@ search_match(int fd, const char *basename, } out[i] = (char)m; } + if (whole && match_body(hdr, hdrlen, len, leaves, nleaves, pool, + out) == -1) { + free(out); + free(hdr); + return (-1); + } free(hdr); *buf_out = out; *len_out = nleaves; blob - 62c0f546829cd6a07af029b8851ddef5fb6c0385 blob + 108e7ceb1611d9fc797ea19eeae6596f60131974 --- src/store.c +++ src/store.c @@ -1496,6 +1496,7 @@ store_exit(void) while ((ss = TAILQ_FIRST(&sessions)) != NULL) store_detach(ss); + pcache_log(); log_debug("store: shutting down"); exit(0); } blob - 7e705ce2ffaab5a8a2440376b5d57698a08f9620 blob + bba0601daaaca5fdca6c7ef25b28c96d2e6f1a03 --- src/store_cmd.c +++ src/store_cmd.c @@ -445,10 +445,8 @@ copy_move_dispatch(struct session *s, const char *tag, return (1); } - if (parse_mailbox_name(&p, mailbox, sizeof(mailbox), &errmsg) == -1) { - session_reply(s, tag, "BAD", errmsg); - return (1); - } + if (parse_mailbox_name(&p, mailbox, sizeof(mailbox), &errmsg) == -1) + return (session_arg_error(s, tag, errmsg)); if (*p != '\0') { session_reply(s, tag, "BAD", "trailing garbage after " "mailbox name"); blob - /dev/null blob + 0694acc1ffb3e92c45054d2d96a511d9ccdbcd5d (mode 644) --- /dev/null +++ src/store_cache.c @@ -0,0 +1,262 @@ +/* $OpenIMAPD$ */ + +/* + * Copyright (c) 2026 David Williams + * + * Permission to use, copy, modify, and distribute this software for any + * purpose with or without fee is hereby granted, provided that the above + * copyright notice and this permission notice appear in all copies. + * + * THE SOFTWARE IS PROVIDED "AS IS" AND THE AUTHOR DISCLAIMS ALL WARRANTIES + * WITH REGARD TO THIS SOFTWARE INCLUDING ALL IMPLIED WARRANTIES OF + * MERCHANTABILITY AND FITNESS. IN NO EVENT SHALL THE AUTHOR BE LIABLE FOR + * ANY SPECIAL, DIRECT, INDIRECT, OR CONSEQUENTIAL DAMAGES OR ANY DAMAGES + * WHATSOEVER RESULTING FROM LOSS OF USE, DATA OR PROFITS, WHETHER IN AN + * ACTION OF CONTRACT, NEGLIGENCE OR OTHER TORTIOUS ACTION, ARISING OUT OF + * OR IN CONNECTION WITH THE USE OR PERFORMANCE OF THIS SOFTWARE. + */ + +/* + * ENVELOPE and BODYSTRUCTURE text, as checked by parser_request(), kept + * per account worker. The key is RFC 9051 SS2.3.1.1's mailbox name, + * UIDVALIDITY and UID; the index's file name is checked on every hit. + * Evicted oldest first, as usr.sbin/smtpd/queue_backend.c evicts. + */ + +#include +#include +#include + +#include +#include +#include +#include +#include +#include + +#include "imapd.h" +#include "log.h" +#include "store_internal.h" + +struct pcache_entry { + RB_ENTRY(pcache_entry) node; + TAILQ_ENTRY(pcache_entry) lru; + const char *mailbox; /* after the struct */ + uint32_t uid; + uint32_t uidvalidity; + const char *basename; /* after the mailbox */ + char *text[PCACHE_NITEMS]; + uint32_t len[PCACHE_NITEMS]; + size_t bytes; +}; + +RB_HEAD(pcache_tree, pcache_entry); +TAILQ_HEAD(pcache_lru, pcache_entry); + +static int pcache_cmp(struct pcache_entry *, struct pcache_entry *); +static struct pcache_entry *pcache_find(const char *, uint32_t, uint32_t, + const char *); +static void pcache_free(struct pcache_entry *); + +RB_PROTOTYPE_STATIC(pcache_tree, pcache_entry, node, pcache_cmp); +RB_GENERATE_STATIC(pcache_tree, pcache_entry, node, pcache_cmp); + +static struct pcache_tree pcache_root = RB_INITIALIZER(&pcache_root); +static struct pcache_lru pcache_age = TAILQ_HEAD_INITIALIZER(pcache_age); +static size_t pcache_bytes, pcache_entries; + +static struct { + unsigned long long hit, miss, add, evict, drop, stale; + size_t peak_bytes, peak_entries; +} pcache_stat; + +/* by name, then UID, so one mailbox's entries are adjacent */ +static int +pcache_cmp(struct pcache_entry *a, struct pcache_entry *b) +{ + int c; + + if ((c = strcmp(a->mailbox, b->mailbox)) != 0) + return (c); + if (a->uid != b->uid) + return (a->uid < b->uid ? -1 : 1); + return (0); +} + +static void +pcache_free(struct pcache_entry *e) +{ + int i; + + RB_REMOVE(pcache_tree, &pcache_root, e); + TAILQ_REMOVE(&pcache_age, e, lru); + pcache_bytes -= e->bytes; + pcache_entries--; + for (i = 0; i < PCACHE_NITEMS; i++) + free(e->text[i]); + free(e); +} + +/* a UIDVALIDITY or file name that differs means another message */ +static struct pcache_entry * +pcache_find(const char *mailbox, uint32_t uidvalidity, uint32_t uid, + const char *basename) +{ + struct pcache_entry key, *e; + + key.mailbox = mailbox; + key.uid = uid; + if ((e = RB_FIND(pcache_tree, &pcache_root, &key)) == NULL) + return (NULL); + if (e->uidvalidity != uidvalidity || + strcmp(e->basename, basename) != 0) { + pcache_stat.stale++; + pcache_free(e); + return (NULL); + } + return (e); +} + +int +pcache_has(const char *mailbox, uint32_t uidvalidity, uint32_t uid, + const char *basename, int item) +{ + struct pcache_entry *e; + + e = pcache_find(mailbox, uidvalidity, uid, basename); + return (e != NULL && e->text[item] != NULL); +} + +/* valid until the next pcache_put() */ +const char * +pcache_get(const char *mailbox, uint32_t uidvalidity, uint32_t uid, + const char *basename, int item, uint32_t *lenp) +{ + struct pcache_entry *e; + + e = pcache_find(mailbox, uidvalidity, uid, basename); + if (e == NULL || e->text[item] == NULL) { + pcache_stat.miss++; + return (NULL); + } + pcache_stat.hit++; + TAILQ_REMOVE(&pcache_age, e, lru); + TAILQ_INSERT_HEAD(&pcache_age, e, lru); + *lenp = e->len[item]; + return (e->text[item]); +} + +/* a failure only leaves the item uncached */ +void +pcache_put(const char *mailbox, uint32_t uidvalidity, uint32_t uid, + const char *basename, int item, const char *text, uint32_t len) +{ + struct pcache_entry *e, *old; + size_t mblen, bnlen; + char *p; + + if (len == 0 || len > PCACHE_BYTES_MAX / 2) + return; + if ((e = pcache_find(mailbox, uidvalidity, uid, basename)) == NULL) { + mblen = strlen(mailbox) + 1; + bnlen = strlen(basename) + 1; + if ((e = calloc(1, sizeof(*e) + mblen + bnlen)) == NULL) { + log_warn("session %u: parse cache entry", session_id); + return; + } + p = (char *)(e + 1); + memcpy(p, mailbox, mblen); + e->mailbox = p; + memcpy(p + mblen, basename, bnlen); + e->basename = p + mblen; + e->uid = uid; + e->uidvalidity = uidvalidity; + e->bytes = sizeof(*e) + mblen + bnlen; + RB_INSERT(pcache_tree, &pcache_root, e); + pcache_bytes += e->bytes; + pcache_entries++; + } else + TAILQ_REMOVE(&pcache_age, e, lru); + TAILQ_INSERT_HEAD(&pcache_age, e, lru); + + if (e->text[item] != NULL) { + pcache_bytes -= e->len[item]; + e->bytes -= e->len[item]; + free(e->text[item]); + e->text[item] = NULL; + } + while (pcache_bytes + len > PCACHE_BYTES_MAX && + (old = TAILQ_LAST(&pcache_age, pcache_lru)) != e) { + pcache_stat.evict++; + pcache_free(old); + } + if ((e->text[item] = malloc(len)) == NULL) { + log_warn("session %u: parse cache text", session_id); + return; + } + memcpy(e->text[item], text, len); + e->len[item] = len; + e->bytes += len; + pcache_bytes += len; + pcache_stat.add++; + if (pcache_bytes > pcache_stat.peak_bytes) + pcache_stat.peak_bytes = pcache_bytes; + if (pcache_entries > pcache_stat.peak_entries) + pcache_stat.peak_entries = pcache_entries; +} + +void +pcache_drop_uid(const char *mailbox, uint32_t uid) +{ + struct pcache_entry key, *e; + + key.mailbox = mailbox; + key.uid = uid; + if ((e = RB_FIND(pcache_tree, &pcache_root, &key)) != NULL) { + pcache_stat.drop++; + pcache_free(e); + } +} + +void +pcache_drop_mailbox(const char *mailbox) +{ + struct pcache_entry key, *e, *next; + + key.mailbox = mailbox; + key.uid = 0; + for (e = RB_NFIND(pcache_tree, &pcache_root, &key); + e != NULL && strcmp(e->mailbox, mailbox) == 0; e = next) { + next = RB_NEXT(pcache_tree, &pcache_root, e); + pcache_stat.drop++; + pcache_free(e); + } +} + +/* a new UIDVALIDITY (RFC 9051 SS2.3.1.1) ends the old entries */ +void +pcache_check_mailbox(const char *mailbox, uint32_t uidvalidity) +{ + struct pcache_entry key, *e; + + key.mailbox = mailbox; + key.uid = 0; + e = RB_NFIND(pcache_tree, &pcache_root, &key); + if (e != NULL && strcmp(e->mailbox, mailbox) == 0 && + e->uidvalidity != uidvalidity) + pcache_drop_mailbox(mailbox); +} + +/* a worker that looked nothing up has nothing to report */ +void +pcache_log(void) +{ + if (pcache_stat.hit == 0 && pcache_stat.miss == 0) + return; + log_info("uid %u: parse cache: %llu hits, %llu misses, " + "%llu added, %llu evicted, %llu dropped, %llu stale, " + "peak %zu entries, %zu bytes", (unsigned int)getuid(), + pcache_stat.hit, pcache_stat.miss, pcache_stat.add, + pcache_stat.evict, pcache_stat.drop, pcache_stat.stale, + pcache_stat.peak_entries, pcache_stat.peak_bytes); +} blob - 09d189f53097622242549e88eafd6fa569c58af5 blob + a37ae3887f2d08276ba8a3bbe156d22bf670f4e9 --- src/store_internal.h +++ src/store_internal.h @@ -153,6 +153,22 @@ struct store_deferred { uint32_t nelts; }; +/* one RFC 5322 mailbox from address_split(); name and route may be NULL */ +struct address { + char *name; + size_t namelen; + char *route; + size_t routelen; + char *mailbox; + size_t mailboxlen; + char *host; + size_t hostlen; +}; + +#define ADDR_MAILBOX 1 +#define ADDR_GROUP 2 +#define ADDR_GROUP_END 3 + struct store_session { uint32_t id; TAILQ_ENTRY(store_session) entry; @@ -229,9 +245,10 @@ int search_match(int, const char *, const struct imsg uint32_t, const char *, char **, uint32_t *); int header_next_field(const char *, size_t, size_t *, const char **, size_t *, const char **, size_t *); -int address_split(const char *, size_t, const char **, size_t *, - const char **, size_t *, const char **, size_t *); -int address_list_next(const char *, size_t, size_t *, const char **, +size_t address_uncomment(char *, size_t); +size_t address_cook(char *, size_t, int); +int address_split(char *, size_t, struct address *); +int address_list_next(char *, size_t, size_t *, int *, char **, size_t *); int parser_ready(void); int locate_message_file(struct cur_snapshot *, int, const char *, @@ -258,9 +275,9 @@ int envbuf_append_str(char *, size_t, size_t *, const int envbuf_append_nstring(char *, size_t, size_t *, const char *, size_t); int envbuf_append_one_address(char *, size_t, size_t *, - const char *, size_t); -int envbuf_append_address_list(char *, size_t, size_t *, - const char *, size_t); + const struct address *); +int envbuf_append_address_list(char *, size_t, size_t *, char *, + size_t); int append_field_nstring(char *, size_t, size_t *, const char *, uint32_t, const char *); int build_envelope(int, const char *, char **, uint32_t *); @@ -292,6 +309,23 @@ int message_body_range(struct cur_snapshot *, int, co #define STORE_INDEX_LINE_MAX 1024 +/* fixed, as usr.sbin/smtpd/queue.c fixes its envelope cache */ +#define PCACHE_BYTES_MAX ((size_t)4 * 1024 * 1024) + +#define PCACHE_ENVELOPE 0 +#define PCACHE_BODYSTRUCTURE 1 +#define PCACHE_NITEMS 2 + +int pcache_has(const char *, uint32_t, uint32_t, const char *, int); +const char *pcache_get(const char *, uint32_t, uint32_t, const char *, + int, uint32_t *); +void pcache_put(const char *, uint32_t, uint32_t, const char *, int, + const char *, uint32_t); +void pcache_drop_uid(const char *, uint32_t); +void pcache_drop_mailbox(const char *); +void pcache_check_mailbox(const char *, uint32_t); +void pcache_log(void); + /* the index is opened after the lock; init with INDEX_LOCK_INIT */ struct index_lock { int lockfd;