Skip to content
Back to talent search
QDQcalibrationquantizationLinuxINT8/INT4evaluation pipelinesaccuracy recoverycomputer visionONNX RuntimePythonONNXDockerLLM fine-tuningLoRA/QLoRAPyTorchmodel compression

Description

AI / ML engineer focused on model optimization, quantization, and deployment. This profile helps teams take models from “works in research” to “runs efficiently in production.” The work includes PyTorch to ONNX export/debugging, INT8/INT4 quantization, calibration, accuracy recovery after compression, latency/memory benchmarking, and fine-tuning/evaluation workflows. Best fit for startups or teams with ML/AI models that are too slow, too expensive to serve, hard to deploy, or losing accuracy after quantization/compression. Interested in freelance, consulting, part-time, or full-time opportunities around model optimization, edge AI, inference efficiency, and applied ML systems. Years of experience are not specified.