NASR(1)Systems Engineer Manual

press ? for keys · nasr(7) for usage

NAME

Nasreddine Hanafi

nasr — systems and backend engineer — Casablanca, Morocco

SYNOPSIS

nasr [] [] [] [] [] [] [] [] [] []

DESCRIPTION

Software engineer at Oracle, working across enterprise backend systems and low-level systems programming.

Most of what I build, I build from scratch. My training at 42 Network was project-based and specification-driven: rather than assembling frameworks, you implement the thing itself — a web server, a shell, a virtual machine, a renderer — in C and C++, from the RFC or the paper up. I would rather understand the layer beneath the abstraction I am using than treat it as a black box.

Professionally that has taken two forms. On Oracle's PGX Vector Database team I wrote C and C++ for GPU-accelerated graph construction, where the constraints are memory layout, bandwidth and parallelism; the work brought index construction to roughly 8–9× faster than a parallel CPU baseline at equivalent search quality. In my current role I build Java microservices for high-availability payment infrastructure, where the constraints are correctness, reversibility and uptime. Both are performance-and-reliability problems; they just live at different altitudes.

My interests sit where systems programming meets AI infrastructure — vector search, GPU computing, and the tooling around them.

OPTIONS

--c, --cppSockets, non-blocking I/O, processes, signals, file descriptors, manual memory management. Debugged with GDB, Valgrind and AddressSanitizer.
--cudaKernel authoring, launch configuration and , GPU memory optimisation. NVIDIA and in production research at Oracle Labs.
--javaHelidon microservices and REST APIs on high-availability payment infrastructure. Maven, Docker. Working toward Oracle Certified Professional, Java SE.
--plsqlStored procedures and queries carrying core business logic against Oracle Database. Schema evolution through , safe and reversible.
--go, --pythonBackend services and FastAPI. Reproducible benchmarking tooling for latency, recall and scalability measurement.
--graphicsRay tracing, shading, bump mapping, ray-surface intersection, computational geometry. Linear algebra implemented by hand.
--webTypeScript, React, Next.js, NestJS, Node.js, Django, WebSockets, PostgreSQL.
--practiceGit, CMake, Makefiles, CI/CD, performance benchmarking, code review, Agile Scrum, Jira, AI-assisted development.

HISTORY

  1. 8f2a1c9
    HEAD → mainOracle — Fullstack Developer

    Enterprise payment systems, where downtime and incorrect state are both unacceptable outcomes.

    • Develop and maintain Java microservices with Helidon exposing REST APIs consumed by multiple internal teams.
    • Write and maintain Oracle PL/SQL stored procedures and queries carrying core business logic and data access.
    • Manage schema evolution through , where every migration must be safe and reversible against long-lived legacy systems in production.
    • Work in an Agile Scrum team across sprint planning, code review, incident response, and cross-functional work with QA, infrastructure and product.

    Java · Helidon · Oracle Database · PL/SQL · REST · Liquibase · Docker · Maven · CI/CD

    2024-08 → present · Casablanca
  2. 3d91b04
    Oracle PGX Team — Research Assistant, Vector Database

    Accelerating approximate nearest-neighbour search. trades a small, measurable amount of recall for a large gain in query speed — but building those graphs is expensive.

    • Implemented HNSW graph construction on GPU in C and C++ using NVIDIA , achieving roughly 8–9× faster index construction than an 8-thread parallel CPU HNSW baseline — measured on (960 dimensions, 1M vectors) and (128 dimensions, 10M vectors) — with equivalent search-time and recall characteristics.
    • Designed around a deliberately hybrid architecture: graph construction on GPU, where the workload is parallel and throughput-bound; ANN search in C on CPU, where traversal is latency-sensitive with irregular memory access.
    • Minimised host-device memory transfers — moving data across is frequently the real bottleneck in GPU pipelines rather than the computation itself.
    • Built reproducible benchmarking tooling in Python measuring latency, recall and scalability across indexing strategies; presented recommendations that fed into the team’s indexing roadmap.

    C · C++ · CUDA · NVIDIA CAGRA · HNSW · Python · Linux · profiling

    2024-02 → 2024-08 · Casablanca
  3. a77e2f5
    1337 School — Software Engineering Track — UM6P, 42 Network

    Peer-to-peer, project-based curriculum with no lectures and no instructors. Progress comes from building projects to specification and defending them in peer review.

    • Systems programming, operating system internals, networking, computer graphics, algorithms and data structures, software architecture.
    • Core projects implemented from scratch in C and C++ on Linux.
    • Trains two things beyond the technical content: reading specifications and documentation directly, and learning independently under time pressure.

    C · C++ · Linux

    2021 → 2024
  4. c04e8b1
    Faculté des Sciences Dhar El Mahraz, Fès — Licence Professionnelle — Mechatronics & Embedded Systems

    Embedded systems, electronics, control systems, sensors, hardware-software integration.

    embedded · electronics · control systems

    2018 → 2021

FILES

~/1337/webserv[+]HTTP/1.1 server written from scratchinstead of nginxHTTP/1.1 server in C++ — POSIX sockets, poll(), CGI, virtual hosts.

Serves many simultaneous clients from a single thread using non-blocking I/O multiplexed with (), rather than spawning a thread per connection. That choice is the architectural centre of the project: threads cost memory and context switches, and they scale badly under high connection counts.

Single-threaded non-blocking I/O introduces its own difficulty. TCP delivers a byte stream, not discrete messages, so a request may arrive split across several reads and a response may only be partially written before the socket refuses more data. Every connection therefore carries its own state machine tracking how much of a request has been parsed and how much of a response remains to be sent.

  • Non-blocking, single-threaded event loop with poll() readiness multiplexing
  • Hand-written HTTP parser: request line, headers, body, handling
  • Per-connection state machines for partial reads and partial writes
  • Virtual host routing by Host header
  • CGI execution via /exec with piped stdin and stdout, and CGI environment setup
  • Configurable routes, custom error pages, request body size limits
  • Correct file descriptor lifecycle management

C++ · POSIX sockets · poll() · HTTP/1.1 · CGI

github.com/Nx21/webserv

DEMONSTRATES — Socket programming, I/O multiplexing, protocol implementation from specification, and the practical realities of stream-based networking.

~/1337/minishell[+]A Unix shell, rebuiltinstead of bashPipelines, redirections, expansion, signals, builtins. C and raw syscalls.

The parsing is the visible part; process orchestration is the substance. A pipeline like cat file | grep foo | wc -l requires forking a process per command, wiring their standard streams together with pipes and , and closing every file descriptor that process doesn't need. Miss one write end and the downstream reader blocks forever, waiting for an EOF that never arrives. Each child must then be reaped with , or it becomes a zombie.

Signals add a second dimension. When a user presses Ctrl-C, the running child should die and the shell should survive — which means handling differently in parent and child, and respecting the constraints on what is safe to do inside a signal handler.

A third subtlety: some builtins cannot run in a child process at all. cd, export and unset mutate shell state, and a child's mutations die with the child — so those execute in the parent.

The command line at the foot of this page is a small implementation of the same ideas: real argument parsing, quoted strings, a pipe operator between builtins, command history — over the content of this site instead of a filesystem. No process model to get right there, but the parsing problem is the same one.

  • Tokeniser and parser handling single quotes, double quotes, escapes, and variable expansion
  • Arbitrary-length pipelines via pipe() and dup2() with strict fd hygiene
  • Input, output, append, and heredoc redirections
  • Process creation and control with fork, , waitpid, including exit status propagation
  • Signal handling that isolates the shell from its children
  • Builtins executed in the parent where shell state must persist

C · POSIX · system calls · process management · signals

github.com/Nx21/minishelltry the shell on this page

DEMONSTRATES — UNIX process model, file descriptor management, signal semantics, and parsing.

~/1337/dslr[+]Logistic regression in C++, accelerated with CUDAinstead of NumPy, PyTorch or cuBLASClassifier from first principles — custom matrix library, hand-written CUDA kernels.

Layered deliberately. A Matrix library handles allocation and linear algebra. A Statistics library handles data analysis and feature scaling. A CUDA layer offloads matrix operations to the GPU. Gradient descent is implemented in multiple variants, and multi-class classification is handled through a one-vs-rest strategy.

Writing matrix operations by hand makes performance considerations concrete in a way library calls don't. Memory layout determines cache behaviour on the CPU and on the GPU. Kernel launch configuration determines . And for small matrices, the cost of moving data across can exceed the compute time the GPU saves — meaning GPU acceleration isn't unconditionally faster, and knowing when not to use it matters.

This project is in active development.

  • Custom modular Matrix library with hand-implemented linear algebra
  • Custom Statistics library for data analysis and feature scaling
  • CUDA kernels for GPU-accelerated matrix operations
  • Multiple gradient descent variants
  • Multi-class classification via one-vs-rest
  • Makefile-based modular build system

C++ · CUDA · GPU computing · linear algebra · machine learning

github.com/Nx21/dslr

DEMONSTRATES — GPU programming, numerical implementation from mathematical definitions, library and API design, and performance reasoning.

~/1337/miniRT[+]A ray tracer built on vector mathinstead of a render engineAnalytic intersection, Phong shading, bump mapping, shadow rays. C.

For each pixel, a ray is cast from the camera through the image plane into the scene and tested analytically against every object — solving the quadratic for sphere intersection, and the corresponding equations for planes and cylinders. The nearest positive root is the visible surface.

Lighting uses the reflection model: an ambient term approximating indirect light, a diffuse term following Lambert's law from the angle between surface normal and light direction, and a specular term producing highlights from the alignment of the reflected ray with the viewing direction. Shadows are determined by casting a secondary ray from each surface point toward each light and testing for occluders — which introduces the classic shadow acne artefact, where surfaces self-intersect at the ray origin, solved by offsetting along the normal by a small epsilon.

Bump mapping adds apparent surface detail without adding geometry. Rather than displacing the surface, the normal is perturbed from a texture before lighting is computed — so a geometrically flat surface responds to light as though it were textured.

  • Analytic ray-object intersection for geometric primitives
  • Phong shading with ambient, diffuse, and specular components
  • Bump mapping via surface normal perturbation
  • Shadow rays with
  • Custom scene description file format and parser
  • Vector and matrix mathematics implemented from scratch

C · computer graphics · linear algebra · ray tracing

github.com/Nx21/miniRT

DEMONSTRATES — Applied linear algebra, computer graphics fundamentals, and translating mathematical models into working code.

~/1337/abstract-vm[+]A stack-based virtual machine and its bytecode languageinstead of a language runtimeCustom bytecode, lexer, parser, polymorphic type system with promotion rules.

The machine is stack-based: instructions push operands, pop them, and push results — the same execution model the JVM uses. The interesting design problem is type safety. Operands may be int8, int16, int32, float or double, are created dynamically at runtime from parsed text, and arithmetic between different types requires well-defined promotion rules. Creating the right concrete type from a runtime string is a factory problem; performing arithmetic across types without losing precision or silently overflowing is a design problem.

Error handling is split by phase. Malformed bytecode is a parse-time error caught before execution begins; stack underflow or division by zero is a runtime error. Each has its own exception hierarchy.

  • Lexer and parser for a custom bytecode language
  • Polymorphic, type-safe operand system with promotion rules
  • Factory-based dynamic operand creation
  • Stack-based execution engine
  • Separated parse-time and runtime error handling with custom exception types

C++ · interpreter design · OOP design patterns · parsing

github.com/Nx21/Abstract-VM

DEMONSTRATES — Interpreter and VM architecture, C++ object-oriented design, design patterns, and language processing.

~/1337/transcendence[+]Real-time multiplayer web applicationServer-authoritative game state over persistent WebSockets. TypeScript, NestJS.

Real-time multiplayer breaks the request-response model that most web applications are built on. The server must push state to clients continuously rather than waiting to be asked, and it must remain authoritative — a client that owns its own game state is a client that can cheat. WebSockets provide the persistent bidirectional channel; the server drives synchronisation on a tick.

  • Real-time bidirectional communication over WebSockets
  • Server-authoritative game state synchronisation
  • Modular NestJS backend architecture with dependency injection
  • User accounts, authentication, and match history
  • Relational data modelling in PostgreSQL
  • React frontend

TypeScript · NestJS · WebSockets · PostgreSQL · React

github.com/Brahim-maaqoul/ft_transcendence

DEMONSTRATES — Real-time systems, fullstack architecture, and stateful backend design.

~/1337/hypertube[+]Video streaming platformProgressive playback via HTTP range requests. Node.js, Django, React.

Streaming media that isn't fully available locally requires serving content through HTTP range requests, so playback can begin from the portion already retrieved while the remainder continues arriving. The platform pairs a Node.js and Django backend with a React frontend and full user management.

  • Progressive video streaming via HTTP range requests
  • Torrent-based media retrieval
  • User registration, authentication, and profile management
  • REST API backend with React frontend
  • PostgreSQL data layer

Node.js · Django · React · PostgreSQL · REST APIs

github.com/Brahim-maaqoul/Hypertube

DEMONSTRATES — Media streaming mechanics, fullstack web development, and multi-service architecture.

ENVIRONMENT

$PWDCasablanca, Morocco — GMT+1
$LANGArabic (native) · English (fluent) · French (intermediate) · German (intermediate, preparing Goethe-Zertifikat B2)
$RELOCATEOpen to relocation — Europe
$INTERESTSsystems programming · backend engineering · GPU computing · AI infrastructure

NOTES

Preparing for the Goethe-Zertifikat B2 German examination — October 2026.

Working toward Oracle Certified Professional, Java SE Developer.

Continuing development on DSLR (see FILES).

EXIT STATUS

0available for new roles
1requires visa sponsorship
2notice period: one month

SEE ALSO