Git Clone

hobby project pure C#

A Git-style version control system built from scratch in C#. Stage files, commit snapshots and manage branches from a command prompt.

Git Clone thumbnail
Language C# / .NET 8
Interface Command prompt
Storage JSON commit snapshots
Size ~1,400 lines of C#

The short versionRebuilding Git's core loop

scope

Stage, commit, branch

Edit files, stage them, commit a snapshot and manage branches, all through a git-style prompt.

approach

Model Git's own data

A commit holds a tree of file and folder nodes, close to Git's own tree and blob objects, and is saved to disk as JSON.

status

Core in place, gaps listed

Staging, commits and branches are implemented. Restoring old commits, merging and remotes are not built yet. But with the core in place they can be added easily.

Deep diveHow it works

From working directory to commit

Working directory WorkingDirectory · files on disk
→
Index added / modified / deleted vs HEAD
→
Commit tree Tree · file and folder nodes
→
Branch commit list + HEAD
→
JSON on disk one file per commit

Each stage is its own class. A commit is built by walking the working directory: staged files are snapshotted, and unchanged files are carried over from the previous commit's tree.

Detecting what changed

Index.cs programming
// Index.GetAllChangedFiles(): compare the working directory with HEAD
// (headFiles = every file in the HEAD commit's tree)
foreach (var workingFile in workingFiles)
{
    if (headFiles.Contains(workingFile.Name))
    {
        // tracked in HEAD
        var fileNode = (FileNode)HEAD.CommitTree.Nodes[workingFile.Name];
        var currentContent = File.ReadAllText(
            WorkingDirectory.GetFullPathFromRepositoryPath(workingFile.RelativePath)
        );
        var d = diff.DiffText(fileNode.Content, currentContent);
        if (d.Length > 0)
            fileChanges[workingFile.Name] = FileChangeStatus.Modified;
    }
    else
    {
        fileChanges[workingFile.Name] = FileChangeStatus.Added;
    }

    // remove from headFiles so leftover = deleted
    headFiles.Remove(workingFile.Name);
}

// anything still in headFiles must have been deleted
foreach (var deleted in headFiles)
    fileChanges[deleted] = FileChangeStatus.Deleted;

Why it matters: every file in HEAD goes into a set and gets checked off as the working directory is scanned, so whatever is left over was deleted. Modified files are found by diffing the stored content against what's on disk.

Commits as trees

A tree, not a flat list. FileNode and FolderNode both implement INode, and folders hold their children, so a commit mirrors the directory structure.

Unchanged files are reused. When a commit is built, files that didn't change are taken from the previous commit's tree instead of being snapshotted again.

Saved as JSON. Each commit tree is written with Newtonsoft.Json into a per-branch history folder and read back when the program starts.

Loading a tree back from JSON

INode.cs programming
// the tree is made of two node types behind one interface
public class NodeConverter : JsonConverter
{
    public override bool CanConvert(Type objectType)
    {
        return typeof(INode).IsAssignableFrom(objectType);
    }

    public override object? ReadJson(JsonReader reader, Type objectType,
                                  object? existingValue, JsonSerializer serializer)
    {
        JObject obj = JObject.Load(reader);
        if (obj["Content"] != null)
        {
            return obj.ToObject<FileNode>(serializer);
        }
        else if (obj["Children"] != null)
        {
            return obj.ToObject<FolderNode>(serializer);
        }

        throw new JsonSerializationException("Unknown INode type.");
    }

    // WriteJson omitted
}

Why it matters: a folder's children can be files or folders, so JSON has to be told which one to rebuild. The converter looks for Content (a file) or Children (a folder).

The command line

git prompt cli
# git <group> <command>   (every group also has "help")
git help

git repo create | edit | remove <file>
git repo list | status
git repo stage <file> | all

git commit <message>
git commit list

git branch create <name>
git branch delete | clone | checkout <id>
git branch list | current

Why it matters: command names live in constant structs (RepositoryCommands, CommitCommands, BranchCommands), shared by the parser and the help text, so a command is renamed in one place.

My partWhat I built

01
systems

Staging & status

Stage files one by one or all at once, and see what was added, modified or deleted against HEAD

02
systems

Commit trees

Snapshots the repository as a tree of file and folder nodes

03
systems

Branches

Create, clone, delete, list and switch branches by ID

04
programming

JSON persistence

Saves commit trees to disk and loads them back with a custom converter

05
programming

Command interface

A git-style prompt with repo, commit and branch commands and built-in help

06
programming

File management

Create, edit and remove files in the working directory from the prompt

Looking aheadNext steps

Check out old commits

Commits are saved, but restoring the working directory from one isn't built yet. It's the first TODO in Repository.cs.

Link commits together

Commits currently sit in a list per branch. Giving each one a parent would make real history, and merging, possible.

Configurable repository path

The folder being tracked is a fixed path in WorkingDirectory, marked with a TODO.