Running A Go Service With systemd

TL;DR
Use Type=notify and send READY=1 over $NOTIFY_SOCKET once you’re actually listening, ping WATCHDOG=1 so systemd can restart you if you hang, wire ExecReload to SIGHUP, then let systemd-analyze security tell you what to lock down.

I’ve written about respecting SIGHUP to reload config and respecting SIGTERM to shut down cleanly. On most Linux machines, the thing sending those signals is systemd. So let’s close the loop and run a Go service under systemd properly.

“Properly” means more than ExecStart= and walking away. systemd can know when your service is actually ready, restart it when it hangs (not just when it crashes), reload it without a restart, collect its logs and sandbox it. Go makes all of that pretty painless, and we don’t need a single dependency to do it. Let’s dive into it.


The Service

We’ll start with the server from the SIGHUP post: one endpoint that returns a message, and a config file we can reload. This time the message lives in a file.

main.go go
package main

import (
	"context"
	"errors"
	"flag"
	"fmt"
	"log"
	"net"
	"net/http"
	"os"
	"os/signal"
	"strings"
	"sync/atomic"
	"syscall"
	"time"
)

func main() {
	configPath := flag.String("config", "/etc/hello/message", "file holding the greeting")
	flag.Parse()
	log.SetFlags(0) // journald adds its own timestamps

	var message atomic.Value
	message.Store("Hello, World!")
	load := func() {
		b, err := os.ReadFile(*configPath)
		if err != nil {
			log.Printf("<4>keeping the old message: %v", err)
			return
		}
		message.Store(strings.TrimSpace(string(b)))
		log.Printf("loaded message from %s", *configPath)
	}
	load()

	http.HandleFunc("/", func(w http.ResponseWriter, r *http.Request) {
		fmt.Fprintln(w, message.Load().(string))
	})

	ln, err := net.Listen("tcp", ":8080")
	if err != nil {
		log.Fatal(err)
	}
	srv := &http.Server{}
	go func() {
		if err := srv.Serve(ln); err != nil && !errors.Is(err, http.ErrServerClosed) {
			log.Fatal(err)
		}
	}()
	log.Println("listening on :8080")

	sigs := make(chan os.Signal, 1)
	signal.Notify(sigs, syscall.SIGHUP, syscall.SIGTERM, os.Interrupt)
	for sig := range sigs {
		if sig == syscall.SIGHUP {
			load()
			continue
		}
		break
	}

	log.Println("shutting down")
	ctx, cancel := context.WithTimeout(context.Background(), 10*time.Second)
	defer cancel()
	if err := srv.Shutdown(ctx); err != nil {
		log.Fatalf("<3>forced shutdown: %v", err)
	}
}

Two small things that matter later. We call net.Listen ourselves instead of ListenAndServe, so we know exactly when the port is open. And a SIGHUP reloads instead of exiting, while SIGTERM falls out of the loop into a graceful shutdown.

Build it and put it somewhere systemd can find it.

shell
➜ GOOS=linux go build -o hello .
➜ sudo install -m 0755 hello /usr/local/bin/hello

The Unit File

Here’s a first version of the unit.

/etc/systemd/system/hello.service ini
[Unit]
Description=Hello, World! as a service
After=network.target

[Service]
ExecStart=/usr/local/bin/hello -config /etc/hello/message
ExecReload=/bin/kill -HUP $MAINPID
Restart=on-failure
TimeoutStopSec=15s
DynamicUser=yes
ConfigurationDirectory=hello

[Install]
WantedBy=multi-user.target

Wait, what’s all that?

  1. ExecReload

    This is what systemctl reload hello runs. $MAINPID is our process, so a reload is just the SIGHUP we already handle.

  2. TimeoutStopSec

    How long systemd waits after SIGTERM before it sends SIGKILL. Our shutdown gives requests 10 seconds, so we give it 15. Keep your timeout inside theirs!

  3. DynamicUser

    systemd makes up a throwaway user for the service every time it starts. We never run as root and never have to useradd.

  4. ConfigurationDirectory

    Tells systemd the service reads config from /etc/hello, and creates that directory if it’s missing.

Add a message, load the unit and start it.

shell
➜ sudo mkdir -p /etc/hello
➜ echo "Hello from systemd!" | sudo tee /etc/hello/message
➜ sudo systemctl daemon-reload
➜ sudo systemctl enable --now hello
➜ curl localhost:8080
Hello from systemd!

Now try the reload. Change the file and ask systemd nicely.

shell
➜ echo "Go To Hell!" | sudo tee /etc/hello/message
➜ sudo systemctl reload hello
➜ curl localhost:8080
Go To Hell!

No restart, no dropped connections. Just like the SIGHUP post, except now systemd is driving.

Logs For Free

Notice we never set up a log file. Anything a service writes to stdout or stderr goes to the journal, which is why we turned off Go’s timestamps with log.SetFlags(0). journald adds its own.

shell
➜ journalctl -u hello -f
Oct 09 09:12:01 server-0 hello[2113]: loaded message from /etc/hello/message
Oct 09 09:12:01 server-0 hello[2113]: listening on :8080
Oct 09 09:13:40 server-0 hello[2113]: loaded message from /etc/hello/message

You might have spotted the <4> and <3> in front of some log lines. That’s a little-known journald feature: a syslog priority at the start of a line sets the log level. <4> is a warning and <3> is an error, so journalctl -u hello -p warning shows only the lines that matter.

Telling systemd We’re Ready

Here’s the problem with our unit so far. systemd considers the service “started” the moment the process exists. If something is waiting on hello (another unit with After=hello.service, say), it might start before our port is open.

The fix is Type=notify. systemd hands the service a socket in the NOTIFY_SOCKET environment variable and waits until the service sends READY=1 over it. The protocol is just text over a Unix datagram socket, so we can write it ourselves.

notify.go go
package main

import (
	"net"
	"os"
)

// notify sends a state change to systemd over the socket in $NOTIFY_SOCKET.
// Outside of systemd the variable isn't set and this does nothing.
func notify(state string) error {
	socket := os.Getenv("NOTIFY_SOCKET")
	if socket == "" {
		return nil
	}
	if socket[0] == '@' {
		// abstract socket: the name starts with a null byte
		socket = "\x00" + socket[1:]
	}
	conn, err := net.DialUnix("unixgram", nil, &net.UnixAddr{Name: socket, Net: "unixgram"})
	if err != nil {
		return err
	}
	defer conn.Close()
	_, err = conn.Write([]byte(state))
	return err
}

That’s the whole thing. Because it does nothing when NOTIFY_SOCKET is empty, go run on your laptop still works exactly the same. Now tell systemd when we’re ready, and when we’re on our way out.

main.go diff
 	go func() {
 		if err := srv.Serve(ln); err != nil && !errors.Is(err, http.ErrServerClosed) {
 			log.Fatal(err)
 		}
 	}()
+	notify("READY=1")
 	log.Println("listening on :8080")
main.go diff
+	notify("STOPPING=1")
 	log.Println("shutting down")
hello.service diff
 [Service]
+Type=notify
 ExecStart=/usr/local/bin/hello -config /etc/hello/message

READY=1 goes out after net.Listen succeeds, so “ready” really means the port is open. If the listen fails, we exit and systemd marks the start as failed instead of pretending everything is fine.

The Watchdog

Restart=on-failure handles crashes. But what about a service that’s still running and completely stuck? A deadlock, a goroutine leak that ate the machine. The process is alive, so systemd thinks everything is fine.

The watchdog fixes that. Set WatchdogSec= and systemd expects to hear WATCHDOG=1 at least that often. If it doesn’t, it kills the service and Restart= brings it back. systemd tells us the interval through WATCHDOG_USEC.

notify.go go
// watchdogInterval returns how often systemd expects to hear from us,
// or 0 if the watchdog is off.
func watchdogInterval() time.Duration {
	if pid := os.Getenv("WATCHDOG_PID"); pid != "" && pid != strconv.Itoa(os.Getpid()) {
		return 0
	}
	usec, err := strconv.Atoi(os.Getenv("WATCHDOG_USEC"))
	if err != nil || usec <= 0 {
		return 0
	}
	return time.Duration(usec) * time.Microsecond
}

That needs two more imports in notify.go.

notify.go diff
 import (
 	"net"
 	"os"
+	"strconv"
+	"time"
 )

Then ping at half the interval, which is what the systemd docs recommend.

main.go diff
 	notify("READY=1")
 	log.Println("listening on :8080")
+
+	if interval := watchdogInterval(); interval > 0 {
+		go func() {
+			ticker := time.NewTicker(interval / 2)
+			defer ticker.Stop()
+			for range ticker.C {
+				notify("WATCHDOG=1")
+			}
+		}()
+	}
hello.service diff
 Restart=on-failure
+WatchdogSec=10s

Before shipping I pointed the binary at a fake NOTIFY_SOCKET with a one-second watchdog to see exactly what systemd would receive:

text
  0.5s  READY=1
  1.0s  WATCHDOG=1
  1.5s  WATCHDOG=1
  ...
  3.6s  STOPPING=1

Ready once, a ping every half second, and a goodbye on SIGTERM. Wow. Much amaze.

Locking It Down

systemd will grade your unit for you.

shell
➜ systemd-analyze security hello
...
→ Overall exposure level for hello.service: 8.2 EXPOSED :-(

Ouch. The good news is that a small web server needs almost nothing from the system, so we can take a lot away. Here’s what I add.

hello.service ini
[Service]
# ...everything from above...
ProtectSystem=strict
ProtectHome=yes
PrivateDevices=yes
ProtectKernelTunables=yes
ProtectKernelModules=yes
ProtectKernelLogs=yes
ProtectControlGroups=yes
ProtectClock=yes
ProtectHostname=yes
ProtectProc=invisible
NoNewPrivileges=yes
CapabilityBoundingSet=
RestrictAddressFamilies=AF_INET AF_INET6 AF_UNIX
RestrictNamespaces=yes
RestrictRealtime=yes
LockPersonality=yes
MemoryDenyWriteExecute=yes
SystemCallArchitectures=native
SystemCallFilter=@system-service
SystemCallFilter=~@privileged @resources
UMask=0077

That reads like a lot, but it boils down to: a read-only filesystem, no home directories, no devices, no kernel knobs, no capabilities, and only the system calls a normal service uses.

shell
→ Overall exposure level for hello.service: 1.2 OK :-)

Add one option at a time and restart in between. If something breaks, journalctl -u hello will usually tell you which one did it.

Conclusion

That’s it! Between the last two posts and this one, our little Go server now reloads on SIGHUP, shuts down cleanly on SIGTERM, tells systemd when it’s ready, gets restarted if it hangs and runs in a pretty tight sandbox, all with the standard library.

If you’d rather not own the notify code, coreos/go-systemd has a daemon package that does the same thing and more. I like knowing it’s just a string over a socket. Happy hacking!