Skip to content

If default sending relay is down/unreachable, every message sent delays #8762

Description

@hpk42

Reported from an operator: they had their relay down for some time, and users who had it as their first relay then got 1-2 minutes delays when they sent a message. It's not a surprise because smtp is currently always tried in the same fixed order, every time. But it's a bug in the user experience, and somewhat of a downgrade from before 2.62 where users could influence what is the sending relay.

Activity

  1. Hocuri commented on Sep 29, 2026

    @Hocuri
    Collaborator

    This problem appeared somewhat quicker than expected, maybe we can agree already on a way to fix it for now? I see three options:

    1. Always choosing the one that was used for sending last
    2. Keeping statistics of relay speed
    3. Starting with the relay that downloaded a message most recently: It was the fastest relay that successfully received this message, so it is probably also the fastest relay that works

    Solution 1. is straightforward, requires a little bit of extra state; 2. requires further discussions, so it's nothing for now; 3. would require no additional state (we already have last_rcvd_timestamp), but IIRC @link2xt was against using receiving behavior for choosing a sending relay since sending and receiving behavior of a server is sometimes different?

    (BTW, good that we didn't choose to try sending relays in random order; otherwise, it would have been way harder for the users to debug the problem, and while some messages would have gone out quickly, it would have affected also users who have this relay, but not as the first one)

  2. link2xt commented on Sep 29, 2026

    @link2xt
    Collaborator

    Always choosing the one that was used for sending last

    I think the simplest solution is to remember the timestamp of last successful sent message for each SMTP transport and sort the transports by this timestamp rather than by rowid. For transports with the same timestamp we can still sort them by rowid, it should not really matter in most cases how they are sorted.

    There is also a somewhat related topic on the forum, just for reference: https://support.delta.chat/t/2-62-broke-mixed-use-of-regular-email-and-chatmail-relays-for-one-identity/5838
    I don't think we want to support the case of sending unencrypted messages mixed with chatmail relays, but there are other messages in the topic and https://support.delta.chat/t/2-62-broke-mixed-use-of-regular-email-and-chatmail-relays-for-one-identity/5838/5 describes a similar problem ("Also, after a few tests (at least in my case), DC keeps trying with the previously set sending relay (set with previous DC version). In case of error the message is not sent and nothing can be done about it. The automatic switching is not happening"), but maybe it just means that the message is sent with the last relay (chatmail) and the message is not encrypted so it is rejected with an error.

  3. self-assigned this
    on Sep 30, 2026
  4. link2xt commented on Sep 30, 2026

    @link2xt
    Collaborator

    I created #8771 to always use the transport that was most recently used. It is good enough to close the issue.


    For something fancier, we have three layers that can be reordered:

    1. Transports
    2. Hosts+ports (what is called "ConnectionCandidate")
    3. IP addresses after DNS resolution if it is a domain name

    Last layer (DNS resolution, TLS, SNI etc.) i would rather leave self-contained, but the first two layers with some refactoring can be merged and then we should be able to interleave "connection candidates" from multiple transports. E.g. try port 993 from the first transport, then port 143 from another transport etc. Otherwise if we remember that some transport was successfully used, when we try it we have to try all "connection candidates" of this transport even if some port always times out because relay operator firewalled some port away. This will still need another table because existing connection_history records successful connection rather than successfully sending a message, so not good enough (worse than #8771) for ordering SMTP candidates, but tables are cheap to add anyway.

  5. duub commented on Oct 2, 2026

    @duub

    I had curiosity of how this process worked and one question I have is if DC choosed the shortest path, like if I want to send a message to someone I have the same relay configured if it would use that relay to send it. Does it have any sense or it's just a human logic that technically doesn't makes sense?

    And another question, if DC checks the connectivity with a relay and knows the status of each one, would be possible to avoid to send messages from the ones that are down?

    Is simplicity in this decision making a priority? Because It would be interesting a more elaborate process like:

    1- Is the relay up?
    2- Do I have configured the relay of the receiver? (if that makes sense, as mentioned before)
    3- Is enough space left in this relay account?
    4- Is the relay able to send this message? (I want to configure my relay to send bigger files than the default setup)

    for example...

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

bugSomething is not working

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions