π― I'm currently working on
Reproducing and extending research papers in DPO/RLHF and RAG β currently a two-stage SFTβDPO pipeline for humor generation, and a GitHub-repo explainer with RAGAS evaluation.
π€ I'm looking to collaborate on
Open-source LLM fine-tuning and evaluation projects β especially anything involving preference optimization, retrieval-augmented generation, or reproducing ML papers on constrained hardware.
π¬ Ask me about
DPO/LoRA fine-tuning on consumer GPUs, reproducing research papers, or anything RAG-related.
β‘ Fun fact
I once found that a published paper's own recommended hyperparameters break the fix it proposes β a 6GB laptop GPU was enough to catch it.
Highlights
- Pro
Pinned Loading
-
qweneq2latx
qweneq2latx PublicFine-tuned Qwen2-VL-7B with 4-bit LoRA to turn equation images into LaTeX, training the whole model on a free Colab T4 with Unsloth and TRL.
Jupyter Notebook 1
-
robonomous
robonomous PublicReal-time multi-object tracking with YOLOv8 detection and BoT-SORT + CLIP ReID β evaluated on MOT17 (43.3% HOTA).
Python 1
-
dataforge
dataforge PublicBuilt a modular Python pipeline for document understanding, from entity extraction and knowledge graph construction to embedding-based retrieval for RAG applications.
Python
-
reddit_mcp_server
reddit_mcp_server PublicBuilt an MCP server that gives LLM agents access to Reddit posts, comments, and subreddit search through PRAW.
Python
-
If the problem persists, check the GitHub status page or contact support.