Nyami
FeaturesPricing
Back to blog
2026-09-20ยท8 min read
native-kernelperformanceloweringv10

How I Improved Python By Removing It

V10 replaces Python's own builtins with C++ implementations in the native kernel, so there is no Python left for a decompiler to reconstruct. Here is what that cost, what it broke, and the benchmark numbers.

Python is too readable. That is not a hot take, it is the reason products written in it are so easy to crack, and for a long time the honest answer to "how do I protect this" was a shrug. The other obfuscators on the market wrapped your code, marshalled it, encrypted it with a key that was sitting right there, and called the result unbreakable. Then someone published a deobfuscator on GitHub and that was that.

So I stopped trying to hide Python. In V10, Nyami removes it.

The idea

It started as a "what if" before V9 shipped: what if every builtin your code calls was not Python's builtin at all, but our own implementation written in C++? Not wrapped, not shadowed. Replaced.

You still write print("hello"). The obfuscator rewrites that call site to a per-build alias, and the alias is meaningless: xyzw("hello") in one build is abc("hello") in the next, and the string argument is encrypted on top of that. If a cracker learns what abc does in one build, that knowledge is worthless in the next one, because the name, the string, and the dispatch table are all generated at obfuscation time.

That is what "lowering" means here. The call site is lowered out of Python and into the native kernel, and what remains at the source layer is an encrypted alias pointing at an implementation that is not in the file at all.

Why the kernel had to exist first

The Native Kernel was introduced before V9 for exactly this. It is where the important data lives, in C++, with C++'s memory management and the security techniques that come with it, while the artifact still executes as Python.

The obvious shortcut would have been Cython. But .pyd files still carry enough Python API surface to reconstruct a great deal, because that API surface is precisely what lets them execute as Python and import as a Python module. I wanted the compatibility and the module interface. I did not want the Python API sitting in the compiled binary. So the lowering had to happen at the Python level, one family at a time.

It was, genuinely, too hard for a niche this small, and I would do it again anyway.

The numbers

Every number below comes from tests/bench/bench_driver.py on CPython 3.12, default build path, the artifact run in its own subprocess with the machine idle, median of 7 runs. Series 1 and Series 2 are two separate builds with fresh artifacts each time, not one run measured twice. A ratio of 1.00x means identical to plain Python, under 1x is faster, over 1x is slower.

FamilySeries 1Series 2
arithmetic / comparison operators1.46x1.56x
not / is / is not / in / not in1.48x1.53x
augmented assignment1.24x1.32x
fusion into single kernel dispatches1.00x0.93x
f-strings and format calls1.70x0.70x
slicing, iter and next1.47x1.41x
str / bytes / bytearray / ord / chr0.19x0.88x
int / float / bool / complex0.81x0.82x
list / tuple / dict / set / frozenset / len / sorted0.94x0.86x
module-service calls0.98x0.99x
protocol-helper calls0.89x0.93x
function / class / iterator helpers1.00x1.00x
import statements0.37x0.37x

Two rows look wrong on purpose. The string row at 0.19x and the format row at 1.70x are not the lowering getting faster or slower, they are the OFF baseline moving, because the default build's per-build decoy content changes how many literals the string-decrypt layer chews through before the timer starts. The ON side was stable across both series: 0.100-0.107s against 0.092s plain for string, 0.574s both times for format. Read those two rows as an unstable baseline, not a 4x win.

The four operator-heavy families sit at 1.24x to 1.56x and that is the floor, not a bug I have not gotten to. Each of those benches fires 11 to 16 operators per loop iteration, and what is left after direct-capturing the pure-compute paths is one Python-level call through a per-build randomized alias, replacing a single CPython opcode. Getting under 1.2x means putting the operator back in source-visible form, which defeats the entire point of the family.

What closed the gap

It did not start here. The first pass was a working mess: format at 6.48x, arithmetic at 3.71x, truth at 2.83x, fusion at 2.21x. Only 3 of 13 families landed at or under 1x. I could not ship that.

The fix was one pass. A large share of the pure-compute lowered calls were still going through a full tamper-guard wrapper on every single call, and that guard was dead code: if tampering is detected the kernel already wipes the key material and kills the process, so there was never a path where the guard could fire and matter. Stripping it off everything that does not touch state or secrets, and leaving it only on I/O, string decrypt, integrity checks, key material and dispatch, dragged the operator families from the 2x-6x range down to 1.2x-1.6x. Fusion and f-strings got dedicated fast paths on top.

That was a real mistake I shipped and then fixed, not a smooth curve.

What it broke

A case study that only shows wins is an ad.

  • Any script using a keyword argument (f(1, c=2), dict(a=1), kwonly args, **kwargs) silently died on Python 3.10 across four of five build paths. An anti-decompiler pass was renaming a function's internal variable names, and Python binds keyword arguments by matching the name against those internal names. Our own fail-closed logic caught the resulting TypeError and killed the process, thinking it was a tamper attempt. It was us breaking our own output on legal Python.
  • An entire anti-decompiler layer was doing nothing on 3.11 through 3.14 because the exception-table validator read a byte format backwards, little-endian instead of the big-endian CPython uses from 3.11 on. When it hit a multi-byte field it could not parse, it quietly reverted the protection layer to unprotected code instead of failing loud. Single-byte cases decoded the same either way, which is exactly why nobody caught it.
  • Zero-argument super() inside obfuscated class methods died, because a function-compression pass compiled each method in isolation and stripped the hidden __class__ reference Python attaches automatically.
  • On the PYTOC path, indexing a list with a constant negative index like lst[-1] could read into adjacent memory and return garbage, because Cython skips bounds checks on constant-looking accesses. The same root cause made min() and max() return the wrong type when mixing ints and floats in one call.
  • Each one got a real fix and a dedicated regression test.

    How I checked myself

    We ran a byte-tamper sweep: 12 single-byte mutations at different offsets across a built artifact, all printable-preserving, on two separate randomized builds. Every run either failed closed with a non-zero exit, or exited 0 with output byte-identical to the untampered version. No run produced partial or wrong output quietly. Eight or nine of the twelve offsets hit the kernel's kill ladder directly (exit 137); the rest died earlier, at the outer decompression layer or as a plain SyntaxError.

    What this still does not cover

  • No allocation-failure testing. Forcing a memory allocation to fail inside the native primitives needs an instrumented build we do not have.
  • No binary-level inspection of the compiled kernel. We scanned the text-level artifact for leaked strings; the .pyd internals are unaudited by us.
  • The outer compression wrapper around the loader has no key on it, so corrupting that specific layer gives a Python traceback instead of a silent kill. It leaks no key or payload, but it does show a little of the loader's structure.
  • The mapped kernel file holds recoverable per-build key material in memory for as long as the process lives. That is an accepted tradeoff of running client-side at all, and I would rather say it than have someone say it for me.
  • Nothing client-side is unbreakable. Someone with enough time, experience and a model will eventually get through, and it will take far longer than with anything else on the market. Do not put credentials in plaintext in your code and assume this saves you.

    Why this matters beyond Nyami

    The idea of removing Python's core while it still behaves as Python sounded like science fiction when I first had it. It turned out to be mostly a grinding retrofit: the family-by-family fixes and the guard-stripping pass landed over about a week.

    Nyami started as a hobby, turned into a passion project, and the next step is turning it into a business that holds itself up. The reverse-engineering community is small and scattered, which is why the entry price is a single EUR 1 token, why the Discord is where it is, and why I want to run CTFs and events there with real prizes.

    If you want to test it, one EUR 1 token runs the full pipeline. If you want to argue with me about any number in this post, the Discord is the place.

    Ready to protect your own Python code?

    Try Nyami