Is Assembly the Answer in Go?

No.
No it's not.

You may leave the blog now if you were looking for the answer. But I do suggest sticking around to see how or why I fell into this rabbit hole, and why the answer might actually be "it depends".

I don't expect you as the reader to have all the knowledge in the world, but I am also not going to discuss what a map looks like in Go.

Part 1: Of Feature Flags and One Engineer's Folly

In our current platform at Futureplay we've successfully avoided making things complex for the last 3 years, more info in this talk here from KubeCon 2026: https://www.youtube.com/watch?v=Aa04SuPhxtA

But as I mention in that talk, and also now here, there's a difference between simple and basic. Yes the idea of keeping it simple and lean still stands but that doesn't mean we are going to avoid functionality, libraries or useful paradigms that are widely used in the industry just for the sake of keeping things "basic". In this case the suspect in question is: feature flags.

We already had 1 or 2 cases of feature flagging but since there were only a few cases setting an env variable that can be managed remotely always did the trick, not this time. Now there are many conflicting opinions on them (as usual within the tech community), some I agree with, some I don't. Although this post is not about feature flags' ups or downs I can still mention the reason why this need arose after 3+ years in production.

Our platform for many years has only served one game: Merge Gardens. Things can be stupid simple when you have only one customer (internal or external). The challenge becomes: what if that number goes up to 2, 3, or 50? What if we want to use all the same functionalities of the platform for many games? What about then? Should we start creating three different auth services within the mono repo which hosts all the services, that are identical except for a few game specific differences for different needs? Or should we stick with one big infrastructure where every game shares the same cluster fleet and create a disaster case where if it blows up the whole game catalogue of the studio for millions of players goes down with it?

Instead we went with the same old mentality, introduce the simplest possible piece of the puzzle to the infrastructure and iterate on it when and where needed. Sounds good right?

Well it could be.. unless you look at some of the feature flag libraries out there —quick disclaimer: by all means, I'm not throwing shade, there are many open source contributors building these brilliant libraries, and it's one of those cases there's something for everybody— I wasn't quite pleased with the complexity they bring with them. Things like full expression languages, dependency graphs, plugin systems, direct network requests (yes each flag lookup is a network request), and more... I basically just wanted remotely configurable if() blocks, that work really really fast with little to no CPU overhead.

I couldn't really find it, so I did the most engineer thing ever and wrote it myself.

Part 2: Nimbool

"Fast feature flag lookups for Go." is the whole title for nimbool and it does what it says here: https://github.com/berkayuckac/nimbool. Just feature flag lookups in a structured way without performance overhead. It has neither bells nor whistles.

Before digging too deep into the architecture and the decisions and methods I've tried, I'd like to clarify why I am so obsessed with the idea of overhead in feature flags.

To those somehow new to servers but stumbled upon this blog, performance has always been, and still is a big concern in the server development world. Mainly due to its one-to-N relation when it comes to a single piece of execution. Meaning every instruction that gets executed, it's not just for one person but it's for millions, tens or hundreds of millions of users, countless billions of times over the years. You might have the best data layouts and the brightest minds trying to recreate A* from scratch just to get 1% better performance, only to erase every benefit it has brought with a simple network call or a wrong string lookup type.

Let's talk some numbers. First we'll dig into the current version of nimbool and compare it to different worst cases, then the other methods that I have tried and failed to see any improvements or good tradeoffs.

BenchmarkEnabled-12                      1000000000               0.4480 ns/op          0 B/op          0 allocs/op
BenchmarkEnabledMiss-12                  1000000000               0.5207 ns/op          0 B/op          0 allocs/op
BenchmarkEnabledNegativeID-12            1000000000               0.2521 ns/op          0 B/op          0 allocs/op
BenchmarkEnabledParallel-12              1000000000               0.3216 ns/op          0 B/op          0 allocs/op
BenchmarkEnabledProcessPercentage-12     1000000000               0.4173 ns/op          0 B/op          0 allocs/op

Using the simplest lookup mode in the library, we're already down to sub nanosecond throughput in this benchmark. Although most likely you'd wanna run something like "deploy the feature -> roll out slowly over time -> enable for all later", which then you would use:

BenchmarkEnabledForPercentage-12          100000000              10.31 ns/op           0 B/op          0 allocs/op
BenchmarkEnabledRandomPercentage-12       234839155               5.115 ns/op          0 B/op          0 allocs/op

In this example we have more instructions due to the complexity of the lookup, we are still at no alloc lookups but the time it takes to complete an operation is significantly more than what it was on the naive one.

It's still pretty fast. It's just a few lines of lookup from the fast Lx caches or memory.

We can also check the cost of each call per year:

Note 1: costs are calculated based off GCP c4d-standard us-central1, vCPU-hour is $0.04725 on-demand
Note 2: CPU time is ns/op × total requests ÷ 10⁹ values in seconds, and thus cost becomes CPU seconds ÷ 3,600 × $0.04725
Note 3: These are CPU-time-equivalent compute costs, actual bills may not directly reflect as savings depending on the various factors such as instances, sizing, utilization etc.

nimbool Enabled0.448 ns
nimbool ForPercentage10.31 ns
Lib A IsEnabled854 ns
Lib B static1,540 ns
Lib B targeting12,818 ns
Lib C EvalFlag29,396 ns
ns/op

Scenario 1: Sustained load, full cost of each call per year

Benchmarkns/opvCPUs @ 1M req/s$/year @ 100k req/s$/year @ 1M req/s$/year @ 10M req/s
Enabled0.44800.00045$0.02$0.19$1.85
EnabledMiss0.52070.00052$0.02$0.22$2.16
EnabledNegativeID0.25210.00025$0.01$0.10$1.04
EnabledParallel0.32160.00032$0.01$0.13$1.33
EnabledProcessPercentage0.41730.00042$0.02$0.17$1.73
EnabledForPercentage10.310.01031$0.43$4.27$42.67
EnabledRandomPercentage5.1150.00512$0.21$2.12$21.17

Scenario 2: Total requests

Benchmarkns/opCPU time @ 1B callsCPU time @ 1T callsCost @ 1T calls
Enabled0.44800.45 s7.5 min$0.006
EnabledMiss0.52070.52 s8.7 min$0.007
EnabledNegativeID0.25210.25 s4.2 min$0.003
EnabledParallel0.32160.32 s5.4 min$0.004
EnabledProcessPercentage0.41730.42 s7.0 min$0.005
EnabledForPercentage10.3110.3 s2.9 h$0.14
EnabledRandomPercentage5.1155.1 s1.4 h$0.07

Scenario 3: In-process evaluation

Library / benchmarkns/opEvals per vCPU per secvCPUs @ 1M evals/s$/year @ 1M evals/s$/year @ 10M evals/s
nimbool Enabled0.4482.2 billion0.00045$0.19$1.85
nimbool EnabledForPercentage10.3197 million0.0103$4.27$42.67
Lib A IsEnabled854.31.17 million0.85$353.59$3,535.90
Lib B static bool flag1,540649,0001.54$637.39$6,373.91
Lib B bool flag with targeting12,81878,00012.8$5,305.24$53,052.42
Lib C EvalFlag29,39634,00029.4$12,166.71$121,667.10

Scenario 4: In-process evaluation, total evaluations

Library / benchmarkns/opCPU time @ 1B evalsCost @ 1B evalsCPU time @ 1T evalsCost @ 1T evals
nimbool Enabled0.4480.45 s$0.0000067.5 min$0.006
nimbool EnabledForPercentage10.3110.3 s$0.00012.9 h$0.14
Lib A IsEnabled854.314.2 min$0.011237 h (9.9 days)$11.21
Lib B static bool flag1,54025.7 min$0.020428 h (17.8 days)$20.21
Lib B bool flag with targeting12,8183.6 h$0.173,561 h (148 days)$168.23
Lib C EvalFlag29,3968.2 h$0.398,166 h (340 days)$385.80

I didn't include network calls because technically network calls need to be investigated more deeply. Meaning yes building and parsing requests, TLS encryption/decryption, HTTP or gRPC, scheduling, garbage collection does cost vCPU but while waiting on network the goroutine can be parked by the runtime poller, thus it doesn't translate directly into equivalent CPU consumption. To still give a rough calculation a network call without taking the in-flight into account would still add tens to hundreds of µs to each lookup, which alone makes it already slower than nimbool. I might come back to this more in depth in a later post.

Numbers are from whatever public benchmark numbers (very few do share, unfortunately) or personal tests I have done on the same/similar machines. Also of course different libraries also perform different amounts of work, some use targeting flags by default or some have different evaluation syntax. To refer back to Part 1: most people don't need most of these functionalities and if I also don't, why should I pay for it?

One caveat here is that the increase might not only be in the CPU cost, there might be network latency per flag check while waiting for the request to fly back with a response. I cannot cover every single angle on code efficiency, but I'm pretty sure you can find good resources on that :) Also as this is not a scientific paper do not expect the most sterile numbers in the comparisons, I'm sure there might've been a song change or two in Spotify while testing these.

Part 3: The Journey

Like with most software, it's good to start with a naive proof of concept even though you know that it's not the thing you want. At first I got it "working":

Level 0: Parse flags into a slice, do a linear scan

Well maybe I lied, this is already a few layers above "just working". True working would be approaches like reading the file and parsing JSON every call or searching the raw JSON for flag string on each lookup. Hopefully I don't have to explain why those are bad options to even begin with in the goal of performance.

func lvl0LinearScan(name string) bool {
	for _, e := range lvlList {
		if e.name == name {
			return e.enabled
		}
	}

	return false
}

Straightforward: we have the list of flags in a slice and just go through this one by one. Using the same payload file benchmark result returns:

BenchmarkLadder/L0_LinearScan-12                  938384              6621 ns/op               0 B/op          0 allocs/op
BenchmarkLadder/L0_LinearScan-12                  903481              6834 ns/op               0 B/op          0 allocs/op

Level 1: Sorted slice, binary search

Next step is we can take the same structure, but this time do a different lookup rather than a naive linear one:

func lvl1BinarySearch(name string) bool {
	i := sort.Search(len(lvlList), func(i int) bool {
		return lvlList[i].name >= name
	})
	return i < len(lvlList) && lvlList[i].name == name && lvlList[i].enabled
}

And the benchmark reflects it with already a huge difference in the op times:

BenchmarkLadder/L1_BinarySearch-12              145803898               41.13 ns/op            0 B/op          0 allocs/op
BenchmarkLadder/L1_BinarySearch-12              149520550               40.43 ns/op            0 B/op          0 allocs/op

Level 2: Map behind an RWMutex

Let's change things up a bit. Rather than a slice why don't we use a hash map for the lookup? This will require a mutex when you have parallel requests flying through. (side note: this can also be a sync.Map to avoid using mutex, though that might make the performance slightly worse depending on how often you update the flags cache)

func lvl2MutexMap(name string) bool {
	lvlMu.RLock()
	v := lvlMap[name]
	lvlMu.RUnlock()
	return v
}

Benchmark, again a really good improvement on this iteration.

BenchmarkLadder/L2_MutexMap-12          537791122               10.70 ns/op            0 B/op          0 allocs/op
BenchmarkLadder/L2_MutexMap-12          545948214               11.05 ns/op            0 B/op          0 allocs/op

Level 3: Map behind an atomic.Pointer

Another take can be to do just one atomic load per read operation and never block, a reload will build a fresh map and swap it without requiring a lock.

var lvlSnap atomic.Pointer[lvlSnapshot]

func lvl3AtomicMap(name string) bool {
	return lvlSnap.Load().byName[name]
}

func lvlReload(m map[string]bool) {
	lvlSnap.Store(&lvlSnapshot{byName: m})
}
BenchmarkLadder/L3_AtomicMap-12                 669088705                8.918 ns/op           0 B/op          0 allocs/op
BenchmarkLadder/L3_AtomicMap-12                 667898660                8.781 ns/op           0 B/op          0 allocs/op

Down to single digit ns per op values now. Things are getting spicy in the diminishing returns curve.

Level 4: Pre-resolved ID lookups (nimbool)

This is where it gets tricky, we are already down to single digit nanosecond operations, to go further down even more we'll need to change our approach to how the library works. Mainly the lookup. String lookup costs too much in the previous versions. Thus we move the cost to cache load or refresh once to save time from millions of lookups per second.

In nimbool you as the developer own the flag IDs. You declare them as constants and thus a lookup becomes a bounds check plus a small read:

"https://github.com/berkayuckac/nimbool/blob/47fe75b18cbe703bfac0afcb0a94ba4b95957aa9/nimbool.go#L50-L60"

const (
	Checkout nimbool.ID = 1
	Search   nimbool.ID = 2
	NewFlow  nimbool.ID = 3
)

// Enabled reports whether a given id is enabled in the active flag payload.
// Percentages between 1 and 99 are decided once per Replace, so this is the
// fastest lookup mode.
func (f *Flags) Enabled(id ID) bool {
	values := f.activeValues.Load()
	i := int(id)
	if values == nil || uint(i) >= uint(len(values.enabled)) {
		return false
	}
	return values.enabled[i] == 1
}
BenchmarkEnabled-12                      1000000000               0.4480 ns/op          0 B/op          0 allocs/op
BenchmarkEnabledMiss-12                  1000000000               0.5207 ns/op          0 B/op          0 allocs/op
BenchmarkEnabledNegativeID-12            1000000000               0.2521 ns/op          0 B/op          0 allocs/op
BenchmarkEnabledParallel-12              1000000000               0.3216 ns/op          0 B/op          0 allocs/op
BenchmarkEnabledProcessPercentage-12     1000000000               0.4173 ns/op          0 B/op          0 allocs/op

At this level CPU has one clock cycle of 0.25ns at ~4 GHz. A 0.45ns op means each loop iteration costs a bit under 2 cycles on average.

A single L1 cache load takes around 3 to 4 cycles and is about 1ns. CPU gets past that because it can issue several instructions per cycle and doesn't wait for one iteration to finish before starting the next. So the core overlaps several of them each at a different stage.

This will hold true as long as the flag array is small enough to stay in L1 cache and the branch predictor learns the pattern. In a real program one flag check in the middle of a request handler might take more, depending on if it hits from L2 or L3 or from the main memory. Depending on the slice size and the architecture's cache line size it might pull flags 0 through 127 in one go, making sequential access fast as it's already in cache. And again depending on the CPU and the manufacturer your L1 L2 L3 cache size, how the CPU is split on the hypervisor and the cloud provider the amount you can have sub nanosecond lookups will also change. Maybe as little as few hundred or as high as thousands flags in L1.

For example the negative ID is even smaller with .25 because the operation never happens, it's just loop overhead.

Level: Assembly

I promise it wasn't clickbait, I have actually tried this method as well to see what else is left on the table.

And to give the same answer as in the beginning, not a lot. Go compiler is already effective at optimizing the code and the instruction sets. And if we were to summarize what we are doing roughly by the instruction:

TEXT ·lvl4IDSliceAsm(SB), NOSPLIT, $0-9
	MOVQ	id+0(FP), AX            // AX = id
	MOVQ	·lvlFlags+8(SB), CX     // CX = len(lvlFlags)
	CMPQ	AX, CX                  // compare id with len
	JCC	miss                        // unsigned id >= len -> miss
	MOVQ	·lvlFlags(SB), DX       // DX = slice data pointer
	MOVBLZX	(DX)(AX*1), BX          // BX = lvlFlags[id], zero-extended
	MOVB	BX, ret+8(FP)           // write the result
	RET
miss:
	MOVB	$0, ret+8(FP)           // return false
	RET

and this is what the Go compiler output is from the nimbool:

        MOVQ    0(AX), AX              // AX = f.activeValues.Load() (plain load, atomic on x86)
        TESTQ   AX, AX                 // check values == nil
        JE      miss                   // nil -> miss (return false)
        CMPQ    0x8(AX), BX            // compare len(values.enabled) with id
        JBE     miss                   // unsigned len <= id (id >= len) -> miss
        MOVQ    0(AX), AX              // AX = values.enabled data pointer
        MOVZX   0(AX)(BX*1), AX        // AX = values.enabled[id]
        CMPL    AL, $0x1               // compare enabled[id] with 1
        SETE    AL                     // AL = (enabled[id] == 1)
        JMP     done                   // skip the miss path
miss:
        XORL    AX, AX                 // return false
done:

Two snippets are almost the same and that is the problem. Compiler has already emitted roughly what one would write by hand. Worst of all, hand rolled version will pay extra costs such as not being able to be inlined as it's an assembly function or Go taking the function through a generated wrapper that moves the argument, both making the call cost more than the work. Would you take the trade off of making it worse and also losing the readability and portability?

With that said there are pretty niche areas you can take true advantage of using SIMD/AVX on the x86 architecture. I found out about this video while I was researching this a while back, give it a watch it's a good one: https://www.youtube.com/watch?v=svDTwAzSxic. We have seen up to double the performance in the same applications when we moved from C3D to C4D without changing anything else, that is AMD EPYC 4th gen Genoa to AMD EPYC 5th gen Turin in non-GCP language. Goes to show that Go's designers, developers, community and contributors are already thinking about these developments just so we can have a great time writing it on the high level!

Part 4: Conclusion

Let's ask the same question again: is assembly necessary to squeeze more performance out of Go?

None of the meaningful performance gains came from touching assembly. We went from linear scan taking thousands of nanoseconds to binary search to maps to immutable snapshots, and to nimbool's current implementation.

The biggest improvement came from changing what work needed to happen in the first place even though the goal stayed the same.

That doesn't mean assembly, SIMD or architecture specific optimization doesn't have a place. Assembly earns its place when there's enough work per call to make up for losing the benefits of Go, or if you know a trick the compiler can't see. You can SIMD over large arrays for checksums, searching, parsing, image or audio processing, or cryptography. If you are already weighing AVX-512 against NEON, you probably didn't need this post anyway.

One thing I haven't covered in this post is actually handling cache distribution. The answer depends on what kind of infrastructure you're running this thing on. Microservices, macroservices, maybe even a monolith? You can pick your poison: a sidecar distributing these caches, message-based incremental cache changes, or different validation and invalidation strategies. None of those choices change the lookup path nimbool optimizes.

There is also the question of readability, and whether a library should be opinionated or non-opinionated. For nimbool I've leaned towards the latter and left the responsibility of managing the flags to the developers themselves.

Hopefully some of these findings, and maybe nimbool itself, are useful the next time you find yourself staring at a hot path and wondering what else is left on the table.

Berkay