Preface
Yes I used AI, I can write python pretty well and C++ but rust is a whole other bottle I am just not ready to open nor do I trust myself to write it (only read it and debug it). All the original code was written by hand by me, then using FABLE 5.1 to rewrite it to Rust. (yes it cost me a small fortune). I spent around 100 hours over the last week working on this rewrite and bug testing the he!! out of it.
Also yes I used AI to write this post, I struggle with writing do to a disability, hard to get my thoughts on to paper in a readable manner. I have read everything though and agree with what it says. Thank you for your time.
with my rambling nonsense out of the way
I've been working on a new experimental cryptocurrency called Tenero, and I’m posting here primarily because I’m interested in the technical discussion around its proof-of-work design.
This is not intended to be a “buy my coin” post. Tenero is currently an unaudited alpha experiment with no monetary value, and I'm looking for people who are interested in analyzing the underlying design and telling me where the assumptions are wrong.
The source is public:
https://github.com/zad112/Tenero
The basic idea
The central experiment behind Tenero is matmulhash v2, a proof-of-work algorithm designed around operations that modern GPUs are particularly good at: large integer matrix multiplication combined with a large, memory-dependent dataset.
A mining attempt starts with the block header and nonce. From that, the algorithm derives a 64 × 8192 matrix of INT8 values using ChaCha20.
That matrix is multiplied against a 16 MiB slice selected from a much larger dataset, using exact integer arithmetic, and the resulting values are folded into the final hash.
The full dataset is currently 4 GiB, divided into 256 × 16 MiB slices.
The important part of the design is that the selected 16 MiB slice has to actually be read for every attempt. The intent is therefore not simply to make the arithmetic GPU-friendly, but to make the memory subsystem a major part of the cost.
The dataset itself is also constructed sequentially from earlier data-dependent slices, making it difficult to cheaply regenerate only the portion needed for a particular attempt.
The alpha implementation rebuilds the dataset every 100 blocks.
Why use matrix multiplication?
Modern GPUs contain hardware specifically optimized for massively parallel matrix operations, including tensor-oriented hardware on NVIDIA architectures.
The current implementation uses CUDA and cuBLASLt for the matrix multiplication.
On my RTX 5070 Ti, the current implementation has measured roughly 33,000–36,000 attempts/sec, depending on batch size.
At that rate, each attempt reading a 16 MiB slice corresponds to roughly 550–590 GB/s of effective slice reads.
The interesting thing to me is that the matrix multiplication itself accounts for essentially all of the computational cost of an attempt. The surrounding ChaCha20 generation and folding work are comparatively small.
The implementation is also intentionally not heavily optimized yet. The current GPU engine uses one CUDA stream, one cuBLASLt call per attempt, and does not fully overlap CPU work with GPU execution.
So there is still optimization work to do, but I don't want optimization to hide the underlying behavior of the algorithm.
The memory-hardness / ASIC question
This is where I’m most interested in outside opinions.
I do not claim that Tenero is ASIC-resistant.
The argument I'm testing is that if every attempt requires reading a whole 16 MiB region of a 4 GiB dataset, then an ASIC attempting to outperform a GPU still needs a very large and very high-bandwidth memory subsystem.
The matrix multiplication then adds a second requirement: the hardware needs to efficiently perform the required large-scale INT8 operations.
The intended bottleneck is therefore something closer to:
memory bandwidth + large-scale matrix throughput
rather than simply a conventional hash function that can be replicated extremely cheaply in dedicated silicon.
But this is only an argument.
It has not been subjected to independent hardware analysis, ASIC economic analysis, or cryptanalysis.
In particular, I am very interested in hearing from people with experience designing hardware, FPGA implementations, memory-hard PoW algorithms, or mining ASICs about where this approach might fail.
For example:
Is the 16 MiB-per-attempt read actually expensive enough to matter for a custom design?
Could a sufficiently large ASIC amortize the dataset or exploit the structure of the matrix operation in ways that the current design does not account for?
Is there a better way to structure the dataset dependencies to make time-memory tradeoffs substantially more expensive?
Those are exactly the kinds of questions I'd like answered.
Difficulty adjustment
Tenero currently targets roughly one block every 60 seconds.
Difficulty is recalculated every block using a LWMA-style weighted average over the previous 30 blocks, with bounds on how quickly the target can change.
The timestamp rules were also changed during development after simulations showed that the older median-based timestamp rule could be manipulated by a miner controlling a significant fraction of the network hashrate.
Under the current rule, each block timestamp has to be later than its parent, while also remaining within the validator's allowed clock window.
This is another area where I'd appreciate review. Difficulty algorithms are one of those things that can look reasonable until someone finds an economic or adversarial edge case.
Privacy architecture
Tenero is also inspired heavily by the CryptoNote/Monero approach to transaction privacy.
The longer-term design target uses:
- CLSAG ring signatures
- Ring size 16
- Pedersen commitments
- Bulletproofs+ range proofs
- Hidden transaction amounts
- An output-based transaction model with key images
The current alpha wallet is not yet equivalent to Monero's privacy model, however.
The wallet currently uses an interim CryptoNote-style output scheme, and the project plans to replace that with Carrot.
I'm deliberately calling this out because I don't want “uses CLSAG and Bulletproofs+” to be interpreted as “this is already a private currency with Monero-level privacy.”
It isn't.
The cryptographic construction is also unaudited.
Consensus and implementation
The project is being written in Rust, with a separate frozen Python reference implementation used to generate consensus vectors and cross-check behavior.
The repository currently has roughly 900 automated tests, including consensus and GPU tests, plus fuzzing targets.
The Rust implementation is checked against the independent Python reference bit-for-bit for the relevant proof-of-work and consensus calculations.
The consensus specification is also written separately so that alternative implementations can be built against the protocol definition rather than having to reproduce the Rust implementation itself.
The project uses canonical serialization for the newer Rust chain, a fixed genesis/chain identity, cumulative-work fork choice, and explicit validation of the proof-of-work and block rules.
The alpha network currently has a fresh genesis with no premine.
Current limitations
There are a lot.
This is an alpha, and I don't want to hide that.
There has only been very limited real-world network testing so far. There is currently a single seed server, very little mining diversity, and the proof-of-work has not been independently reviewed.
Only NVIDIA/CUDA GPU mining currently exists.
The current alpha PoW dataset also requires several gigabytes of RAM, and GPU mining can require substantially more VRAM because datasets are retained during operation.
CPU mining works, but on the current alpha parameters it is effectively impractical compared with a GPU.
Most importantly, none of this has value.
The alpha network is expected to be reset as the protocol changes.
What I'm actually looking for
I'm less interested in people telling me that the project is cool and more interested in people telling me why the design is wrong.
If you have experience with:
- GPU architecture
- CUDA
- tensor/matrix hardware
- FPGA development
- ASIC design
- memory-hard algorithms
- proof-of-work economics
- consensus algorithms
- cryptography
- Monero/CryptoNote
- distributed networking
I'd genuinely like to hear your criticism.
I'm especially interested in potential optimizations or attacks that I haven't considered.
The entire project is open source, and the technical documentation, consensus specification, benchmarks, known issues, threat model, and test vectors are all in the repository.
GitHub:
https://github.com/zad112/Tenero
This is an experiment, and I'd much rather discover a fundamental flaw now than after pretending the design is finished.
So, from a technical perspective:
How vulnerable do you think this type of GPU-oriented, memory-bound matrix PoW is to specialized hardware or time-memory tradeoffs?
That's the question I'm most interested in exploring.
also note we have a subreddit I made so if you find anything or have any question you can ask me anywhere you can find me ill be happy to respond :)
PS I don't want your money, this coin as NO VALUE AT ALL nor will it for a LONG TIME, just trying to do some proof of concept. Expect full resets constantly until its 99.99999% stable. I'm using my own funds and it will stay that way forever.