Methodology

What a Majken report is built from, what it can't see, and how to use it fairly.

What we measure

Everything comes from public GitHub data at the time of the analysis, unless the candidate connected their own repositories (see verified reports).

Their own repositories
Public repositories the candidate owns, ranked by recent activity. Forks, dotfiles and config repositories are left out.
Merged pull requests
Pull requests the candidate opened in other people's projects that were reviewed and merged. Someone else decided that work was good enough to accept. Counts are GitHub search totals and leave out repositories on their own account and in organisations they are a public member of (as many organisations as fit in one search, those with most of their pull requests first).
Code reviews
How many pull requests by other people the candidate reviewed (a GitHub search total). How they handle requested changes comes from a sample of recent pull requests, and is only shown with at least 5 reviewed pull requests; its subscore needs at least 4 merged pull requests where changes were requested, and only merged pull requests can give a negative example. The length of review comments isn't scored: it says nothing about their quality.
Languages
Each language's share of the code in their own repositories, from GitHub's byte counts. It shows what they write in, not how well.
Frameworks
How many of their repositories declare each framework in a manifest file: package.json, requirements.txt, pyproject.toml, go.mod, Cargo.toml or Gemfile. A count, not a skill level.
AI review of up to 4 pieces of code
Only code the candidate wrote. In their own repositories, a file is reviewed only when git blame shows they wrote at least half of its lines and at least 40 of them (25 when they wrote at least 90% of the file, a whole small module), and the reviewer is told which lines are theirs. Lines from a repository's first commit, or from a commit that adds more than 1,000 lines or touches more than 15 files, don't show who wrote them. The same goes for files that are entirely from the first commit of a repository with at most 2 commits, when the repository is theirs and they made most of its commits; otherwise such files are skipped, and so are commits that say they came from a template or a project generator. When a file is only theirs through such a commit, GitHub code search checks it first: if an older repository of someone else has the same code, it's skipped; if nothing is found, it's reviewed and marked “Imported in one commit; no copy found elsewhere”; if the search can't run, files with lines from later commits are reviewed instead and it's skipped. In coursework repositories such files are always skipped. From projects outside their own account, only the diff of a merged pull request with at least 20 added lines of code is reviewed (blank lines and comments don't count, stylesheet lines count half, and files it removes are left out), never the whole file. We size up to 50 of their recent merged pull requests, spread across projects, and check the largest first by added lines of code (not tests, docs, lockfiles, generated or vendored files), with a mild preference for recent ones. Pull requests into organisations they are a public member of, such as an employer, count and are labelled as such. Up to 2 of the 4 slots go to such pull requests. Starter templates, generated files (including a Rails schema dump and migrations an engine copies in), vendored, minified and copied files (with another project's license or attribution header, or an @author or $Id: tag naming someone else) are skipped, and so are bundled libraries, new files over 500 lines, and folders a pull request adds a LICENSE, COPYING or README to (vendored code). A pull request adding more than 1,000 lines of code or touching more than 15 code files is a bulk change: only its files with at most 150 added lines are reviewed. Before a public file or pull request is reviewed, GitHub code search looks for two of its distinctive lines (up to 6 searches per report, files imported in one commit first, at most 10 a minute across all reports, results kept for 30 days); if at most 3 other repositories have both and one of those files appeared before the candidate's file or pull request, it is skipped as code published elsewhere. Matches in documentation don't count, and lines found in more than 3 repositories are treated as common boilerplate, not a copy. When the search can't run because of GitHub's rate limit, any other file or pull request is reviewed and marked “copy check skipped”. Code that was adapted and then edited can't always be found this way, so a reviewed file that is mostly the candidate's own may still contain a block adapted from public code. Test code is left out of the reviewed excerpt unless there is nothing else, and then the reviewer is told it is reading tests. Each review scores against a fixed 1 to 10 rubric with worked examples, and every strength and weakness must cite the file and line numbers it is about; a review may list no weaknesses when there is none worth raising. The model is from OpenAI.
Interview questions
Questions about specific lines that were reviewed, in their own files or in a pull request's changes (named by file and line), each with a note on what a good answer covers. Never questions about something they haven't done.

What we don't see

  • Private repositories, including most work done for employers.
  • Code on GitLab, Bitbucket, self-hosted Git or company GitHub organisations we can't see.
  • Anything outside code: product sense, mentoring, how someone works in a team day to day.
  • Whether a candidate wrote code by hand or with an AI assistant.

Many strong developers have almost no public code. A thin report tells you to look elsewhere for evidence; it says nothing bad about the person.

AI-written code

We don't try to detect AI-generated code. No tool does this reliably, and most developers now use AI assistants anyway.

Instead, the report puts weight on evidence an AI can't fake on its own: pull requests other maintainers reviewed and merged, reviews the candidate gave, and code in repositories they own and maintain. The interview questions then let you check that the candidate understands what they shipped. Someone who made the decisions in that code can explain them.

How confidence works

Every report has a confidence level based on how much of the candidate's own work we saw: the lines they wrote in the reviewed code, and their merged pull requests to other people's projects. Each report says what it rests on, for example “Based on 4 files, 830 lines they wrote, 12 merged PRs to other projects”. The overall score starts from the code score: the review scores, each weighted by the square root of the lines they wrote in it, so one large file can't outweigh several reviewed pull requests. A communication score only counts when two parts are measured: team contributions (at least 5 merged pull requests to other projects or 10 reviews given) and documentation (at least 5 pull request descriptions or comments). When it is higher than the code score, it raises the overall score by 40% of the difference; it never lowers it, because missing public collaboration isn't evidence against someone. With low confidence there is no overall score, and no code, communication or critique score either. Scores are shown as the one-point range that contains them, such as 7 to 8 for 7.6, because they aren't precise enough for decimals. When GitHub doesn't answer in time, the report is finished with what it has, missing stats are shown as unavailable, and it doesn't use up a free report.

High confidence
At least 600 reviewed lines they wrote across at least 2 repositories, or at least 400 lines plus at least 10 merged pull requests to other people's projects.
Medium confidence
Everything between low and high. Read the scores as a rough indication. Someone with at least 25 merged pull requests outside their own account (organisations included) is at least medium as soon as one file or pull request of theirs could be reviewed, even if it is under 150 lines.
Low confidence
Fewer than 3 merged pull requests to other people's projects, and either under 150 reviewed lines they wrote or code from fewer than 2 repositories; or no reviewed lines that count as theirs at all. Lines in coursework repositories don't count, since they may be starter code. A repository counts as coursework from clear signs in its name or README, such as “homework”, “bootcamp” or “Lab 3”, and never when it has 50 or more stars. There's not enough of their own code to judge, so we don't give an overall score. Ask for a code sample or a take-home task instead. When someone has at least 25 merged pull requests outside their own account but none of them, and no file of theirs, was large enough to review, the report says exactly that and shows their pull request totals instead.

People make the decision

A Majken report supports a hiring decision made by people. It should never be the only basis for rejecting someone. Scores are an AI-assisted reading of public data and can be wrong; they aren't objective measurements of ability.

We recommend telling candidates that you look at public GitHub work, and giving them the chance to add context, such as pointing you to private work or explaining a gap.

Questions about how a report was made? Get in touch or read the FAQ.