Skip to content

Hostname changing ip management - #225

Draft
bisegni wants to merge 12 commits into
epics-base:masterfrom
bisegni:feature/dns-hostname-reconnection
Draft

bisegni wants to merge 12 commits into
epics-base:masterfrom
bisegni:feature/dns-hostname-reconnection

Conversation

@bisegni

@bisegni bisegni commented Sep 23, 2026

Copy link
Copy Markdown
Contributor

PVXS Hostname Re-resolution and Name Server Hostname Support

Summary

This change allows PVXS clients to keep working when a configured hostname starts resolving to a different IP address. It also allows EPICS_PVA_NAME_SERVERS to contain a hostname instead of requiring a literal IP address.

The client periodically resolves hostname entries again. When the resolved IP changes, PVXS uses the new address for searches or name-server connections without requiring a client restart.

Scope

This change covers the following configuration variables:

  • EPICS_PVA_ADDR_LIST accepts literal IP addresses and hostnames. Hostname entries are retained after startup and periodically resolved again for UDP searches.
  • EPICS_PVA_NAME_SERVERS accepts literal IP addresses and hostnames. Hostname entries are periodically resolved again for TCP connections.

Literal IP addresses keep their existing behavior and are not resolved again. Hostname entries are checked approximately every 10 seconds. For EPICS_PVA_ADDR_LIST, the initial address and the original hostname are both kept. If the hostname later resolves to a different IP, the search destination is updated and subsequent search retries use the new address. For EPICS_PVA_NAME_SERVERS, a changed address causes the old connection to be replaced with a connection to the new address. Disconnected name-server connections are also retried during this check.

Previous Behavior

EPICS_PVA_NAME_SERVERS accepted only literal IP addresses. A hostname in this variable could not be configured successfully.

Hostname entries in the supported client configuration were resolved when the client started and kept using that address. If the hostname later resolved to a different IP, PVXS continued sending searches or connection attempts to the old address.

Configuration Examples

Both variables can use a literal IP address or a hostname:

EPICS_PVA_ADDR_LIST=ioc-pvxs-service.epics-test.svc.cluster.local:5075
EPICS_PVA_NAME_SERVERS=ioc-pvxs-ns-service.epics-test.svc.cluster.local:5075

The hostname must reach PVXS unchanged. A wrapper or container entrypoint that converts it to a literal IP before starting PVXS prevents periodic re-resolution.

Reproduce the Issue

Use the test repository epics-k8s-reconnection-test to run the test in a local Kubernetes cluster.

Prerequisites

Install the following tools before starting:

  • Git, to clone the test repository
  • Docker 20 or newer, with the Docker daemon running
  • kind 0.20 or newer, to create the local Kubernetes cluster
  • kubectl 1.27 or newer, to inspect and manage the cluster
  • Helm 3.x, to install Cilium

The test repository checks for kind, kubectl, docker, and helm when setup-cilium.sh starts.

git clone https://github.com/slaclab/epics-k8s-reconnection-test.git
cd epics-k8s-reconnection-test

Build and start the Cilium cluster. To reproduce the issue with the old behavior, use the repository defaults:

./setup-cilium.sh up

it should print something like:

[INFO]  PVXS_REPO=https://github.com/epics-base/pvxs.git  PVXS_BRANCH=1.5.2
[INFO]  Creating kind cluster 'epics-test-cilium'
Creating cluster "epics-test-cilium" ...
 ✓ Ensuring node image (kindest/node:v1.37.0) 🖼️
 ✓ Preparing nodes 📦 📦 📦
 ✓ Writing configuration 📜
 ✓ Starting control-plane 🕹️
 ✓ Installing StorageClass 💾
...

the script will start building epics and pvxs using the PVXS_REPO and PVXS_BRANCH. At the end verify that all monitor pods are receiving updates before changing the services:

kubectl -n epics-test get pods
kubectl -n epics-test logs -f \
  -l 'app in (monitor-ca,monitor-pva,monitor-pvxs,monitor-pvxs-ns)' \
  --prefix=true

In another terminal, delete and recreate the services with the repository helper. This assigns new IP addresses to the services:

./recreate-services.sh
kubectl -n epics-test get svc

After the services have new IP addresses, restart both PVXS IOC deployments:

kubectl -n epics-test rollout restart \
  deployment/ioc-pvxs deployment/ioc-pvxs-ns
kubectl -n epics-test rollout status deployment/ioc-pvxs
kubectl -n epics-test rollout status deployment/ioc-pvxs-ns

With the old PVXS behavior, monitor-pvxs and monitor-pvxs-ns stop receiving updates because they continue using the addresses resolved before the service recreation.

To verify the fix, rebuild the test repository with the patched PVXS repository and branch:

PVXS_REPO=https://github.com/<owner>/<patched-pvxs-repository>.git \
PVXS_BRANCH=<patched-branch> \
./setup-cilium.sh up

Repeat the service recreation and IOC restart steps. Both PVXS monitors should resume receiving updates after the hostname entries are checked again.

To clean test environment:

./setup-cilium.sh down

Replace AsyncResolver (async DNS on separate thread) with synchronous
setAddress() calls, matching the startup code path. Hostnames are now
preserved in Config and Channel so periodic re-resolution and reconnect
can re-resolve using the same sync mechanism used at initialization.

- Add isHostname() utility to detect hostname vs IP addresses
- Preserve original hostnames in Config for addressList and nameServers
- Store forcedServerHostname on Channel for reconnect re-resolution
- Add periodic DNS recheck timer for search destinations
- Remove AsyncResolver, clientresolver.cpp/h, and dedicated DNS thread
Extend onDNSRecheck() to also re-resolve nameserver hostnames, not just
search destinations. If the IP changes, the old connection is cleaned up
and a new one is established immediately.
startNS() bailed early when nameServers was empty, leaving
dnsRecheckTimer dormant. This made onDNSRecheck()'s searchDest
re-resolution path dead code — hostnames from EPICS_PVA_ADDR_LIST
were resolved once at startup but never re-checked.

Now startNS() also checks whether any searchDest entry carries a
hostname and starts the DNS recheck timer in that case. The TCP
nameserver reconnect timer (nsChecker) remains gated on
nameServers being non-empty.
Connection::build() inserts the new connection into connByAddr keyed by
address. Assigning the result to ns.conn then dropped the last reference
to the old connection, whose destructor runs cleanup() and erases
connByAddr[peerAddr]. For a stable address (e.g. a Kubernetes ClusterIP
that is unchanged across pod restarts), this erased the freshly inserted
entry, leaving search replies unable to find the nameserver connection.

Reset ns.conn before build() so the old destructor's erase runs while
the dead entry is still mapped, letting the fresh insert survive. Applied
to both onNSCheck() and onDNSRecheck().
…name servers

self.tcp_port==0 fallback checked self.nameServers before it was
populated by EPICS_PVA_NAME_SERVERS parsing, so it never fired; a
bare hostname without an explicit port got resolved and cached with
port 0, breaking later DNS re-resolution keying.
…change

onNSCheck was reconnecting hostname-based nameservers to potentially stale
IPs (OS DNS cache) every 10s, while onDNSRecheck only rebuilt the connection
when the resolved IP changed. When a k8s service is deleted+recreated with a
new IP, DNS may still return the old IP during propagation — so onDNSRecheck
never triggered a reconnect, and onNSCheck kept hammering the dead old IP.

Fix by:
- onNSCheck skips hostname entries entirely; onDNSRecheck owns them
- onDNSRecheck reconnects when IP changed OR conn is Disconnected, so a
  downed connection always retries even if DNS hasn't propagated yet
- Reduce dnsRecheckInterval 30s -> 10s to match tcpNSCheckInterval and
  recover within typical k8s DNS propagation time
Raises the connect/reconnect log lines in startNS/onNSCheck/onDNSRecheck
from debug to info, and includes the hostname on connect, so nameserver
contact attempts are visible without enabling debug logging.
Keeps the added hostname text in the log messages but reverts the
level back to debug, undoing the info bump from the prior commit.
Debugging why EPICS_PVA_NAME_SERVERS hostname association is lost at
runtime (reconnects go through the bare-IP onNSCheck path instead of
onDNSRecheck). Traces hostnameMap population in split_addr_into and
the lookup in ContextImpl's nameServers build loop.
… resolution

setAddress() copied the pre-colon substring into a fixed
scratch[INET6_ADDRSTRLEN+1] (47-byte) buffer and threw "IPv4 address
too long" whenever it overflowed, before ever reaching the
evutil_inet_pton/GetAddrInfo fallback used for hostname resolution.
Kubernetes Service FQDNs (e.g. foo-service.namespace.svc.cluster.local)
routinely exceed 46 characters, so any such hostname given via
EPICS_PVA_NAME_SERVERS was rejected outright instead of being resolved,
silently defeating the DNS-recheck/hostname-tracking reconnect logic.
Use a std::string instead of a fixed buffer to remove the length limit.
epicsEnvUnset() is only available on newer EPICS base (7.0+) and broke
the build on the "Native Linux with 3.14" CI job. Add a small
unsetEnv() helper using unsetenv()/_putenv() instead.
@bisegni

bisegni commented Sep 28, 2026

Copy link
Copy Markdown
Contributor Author

@mdavidsaver what do you think about this behaviour?

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant