I’ll just provide my own example: my homelab consists of 6 Kubernetes nodes placed across the country. Some differ by ISP, some are placed in different cities, one is hosted on a cloud provider. Basically it’s a very cheap variant of geo-replicating my workloads.

Two of these nodes are visible from the Internet and have a static IP address; one node also has an IPv6 address. Each node hosts an authoritative DNS server (CoreDNS) for my personal domain pootis.network; and the .network TLD has glue records which point to IPs of these two nodes. This is a classic “self-hosted DNS” scenario.

Here’s an excerpt from my zonefile so you can understand the setup better:

$ORIGIN pootis.network.
$TTL 300

@       SOA     ns1.pootis.network. admin.pootis.network. (
  2026082001
  1200
  300
  1209600
  300
)

; Nameservers and glue records
@       NS      ns1.pootis.network.
@       NS      ns2.pootis.network.
ns1     A       178.44.116.85
ns2     A       91.219.150.30
ns2     AAAA    2a06:dd00:1:4::4189

This 5-record block (NS/A/AAAA) is mirrored into the .network zone by my domain registrar (plus DS for DNSSEC but that’s another thing).

As such, my DNS becomes fully independent - and, in theory, if one of my externally-facing nodes breaks, let’s say ns1, then DNS resolvers all over the world (forwarders, recursive, and such) will fall back to ns2, and everything will keep working. Kubernetes will also reorganize the pod placement so all my workloads are available again after a slight downtime.

That would have been great, if it worked as described, but apparently, after one nameserver in my zone fails, then the resolvers… just give up? Let’s say ns1 failed but ns2 is working. The parent zone still points to both nameservers. My external resource records (websites and other stuff) at this point would have already been auto-reconfigured by a custom k8s controller to point to the IP addresses of the node that hosts ns2. Simplifying: the entire world basically sees this after ns1 fails and after TTL caches expire:

; all of this has very low TTL, 5 minutes or so

@       NS      ns1.pootis.network. ; from .network 
@       NS      ns2.pootis.network. ; from .network

ns1     A       178.44.116.85 ; broken. Either from .network glue or from my auth DNS
ns2     A       91.219.150.30 ; either from .network glue or from my auth DNS
ns2     AAAA    2a06:dd00:1:4::4189 ; same

; my-website     A       178.44.116.85 ; does not appear because ns1 is broken- my LB already removed it from the set
my-website     A       91.219.150.30 ; fronted by a pair of CNAMEs due to loadbalancing but still
my-website     AAAA    2a06:dd00:1:4::4189 ; same

But even if I query 1.1.1.1 directly for my-website’s record, it just doesn’t work most of the time because the resolver pins itself to ns1 which is currently failing, or it selects ns1 and does not even care to try ns2.

To be precise: some resolver implementations DO fall back to ns2 as expected, but most of them just pin themselves to ns1 and then outright refuse to resolve the records in my zone.

And there’s actually no reasonable way out, as far as I can see:

  • moving my DNS infra somewhere else (CloudFlare, for example) is unacceptable since I would like for my homelab to be as independent as practically possible;
  • anycasting, or running a fully-fledged BGP AS is also impossible because that costs a lot of money and I’d like for my homelab to fit into a $10/month budget with room to spare;
  • “live-patching” the NS and glue records in the parent zone (.network), to keep up with the set of my working nodes, is possible, but very unwieldy and somewhat hard to accomplish.

There’s a lot of custom machinery that keeps my workloads running and accessible after a node failure, but all of this becomes completely moot when authoritative DNS is the bottleneck.

Has anyone been running a similar stack and encountered this problem? I’m aware that the answer is usually “host your DNS at CloudFlare” or “use the registrar’s DNS infra” but still…

  • farcaller@fstab.sh
    link
    fedilink
    English
    arrow-up
    2
    ·
    1 day ago

    BGP anycast person here. If you have any presence in the region RIPE operates in then the pricing is within the homelab reach. ASN and a block of /48 would be about 70 EUR/year.

    Alternatively, something like route64 would happily tunnel you IPs they announce for about 2 EUR/month.

    • buedi@feddit.org
      link
      fedilink
      English
      arrow-up
      1
      ·
      19 hours ago

      May I ask how you got this started? Hosting your own ASN, as far as I understand only works if my ISP would actually route traffic to my ASN, right? I am thinking about getting into IPv6 for self-hosting and I could ask my ISP to change my current setup (I only have IPv4, but a public one, without CGNAT) and I do not trust them that they mess this up. So my preferred way would be to not touch anything on the ISP side and host my own ASN and find a Sponsor for a /48 Block. I still cannot wrap my head around this.

      Just that I understand you correctly: You got your ASN and /48 block from RIPE (or a Sponsor I assume) and you host your own AS? Or is the AS hosted by someone else? If the latter, I wonder how traffic can find to your home or to your Server locations.

      Nothing of this would work without getting in touch with my ISP, and I fear the usual resedential IPs will not care.

      • farcaller@fstab.sh
        link
        fedilink
        English
        arrow-up
        1
        ·
        16 hours ago

        First on how to get an ASN: you can buy it for reasonably cheap from a LIR. Some will even toss a free /48 with that. Happy to offer names in private so that there’s no advertising. Expect a budget quoted above.

        Once you have an ASN, you need to get an upstream - actually two as RIPE mandates at least two (otherwise why’d you need an ASN). Some LIRs would offer transit with ASN purchase. You can upstream via your ISP, if they allow you to (that’s very rare). Another option is a tunnel (there are free and paid ones) or a VM somewhere (some cloud providers offer to set up bgp with VMs they host). Generally, free ones are enough for basic stuff. Not much bandwidth and oftentimes IPv6 only, but you don’t pay anything either. Besides, you can ask around in various network related chats. Practically, I can offer ip transit with some marginally low burstable bandwidth, and that’s pretty common. You can look/ask around https://discord.gg/ipv6 for example.

        For getting ASN to your homelab you’re looking at a tunnel option, most probably. Great if you have static ipv4 - allows you to use more common tunnels, but is still doable with a floating IP (e.g. check bgptunnel).

        • buedi@feddit.org
          link
          fedilink
          English
          arrow-up
          1
          ·
          4 hours ago

          Thank you very much for your reply. So the bottomline is, that when bound to a residential ISP, we most likely have no chance to utilize our own AS and IPv6 Prefix without the help of a 3rd party (ASN peering aside, which is needed of course, no matter what). So it’s usually involving tunnels to get it up and running. That solves the mystery somewhat for me and I think I know in which direction I have to do my research.

          I think I will start prototyping with route64 (which you mentioned above), but if I get this up and running, I would appreciate it if you could PM me the LIRs you have on your list. If I can wrap my head around this, if the fees are still in the ballpark you mentioned above, that sounds very reasonable to use this in my Homelab and gather experience with IPv6 and AS.

          Thanks again for your help :-)

    • Dave@lemmy.pootis.networkOP
      link
      fedilink
      English
      arrow-up
      1
      ·
      1 day ago

      That does seem to be a good solution, thank you for the recommendation!

      Running an AS and obtaining a /48 through a sponsoring LIR seems to cost about $150/year in my country, but hardly anyone (except hosting providers and large companies) does this, since self-hosting, especially more complicated stuff, isn’t really that popular here; people are mostly uninformed that it even exists.

      But route64 is apparently completely free (donations are welcome); they provide a /56 PA-like IPv6 block carved out of their PI, and they also handle BGP stuff, and I don’t even have to pay for any of this - overall, a great choice, considering my constraints. Anycast to multiple tunnels costs money (maybe that’s what you meant with 2 EUR/month?) but that’s pretty much lunch money so it would be OK with me.