This changes the stdout/stderr buffering behavior of Python. Without it,
indexing scripts don't stream updates and use really big buffers.
Signed-off-by: Théo Lebrun <theo.lebrun@bootlin.com>
Instead, start from $0 and move back up two times. So, something like:
./elixir/utils/index
./elixir/utils
./elixir
./elixir/update.py
Signed-off-by: Théo Lebrun <theo.lebrun@bootlin.com>
Stop writing a global file when initializing projects. This can cause
permission issues. We instead pass the option manually for each Git
process call using:
git -c safe.directory=...
Signed-off-by: Théo Lebrun <theo.lebrun@bootlin.com>
Previously, to start an indexing from scratch:
./utils/index /srv/elixir-data musl https://git.musl-libc.org/git/musl
This is annoying as the script already has the remote URLs for all known
projects. Now, a call without remote will automatically add the remote
URLs matching the project name:
./utils/index /srv/elixir-data musl
This copies the behavior that was previously only implemented for --all.
Signed-off-by: Théo Lebrun <theo.lebrun@bootlin.com>
New script utils/index does an automatic call to `git gc --auto` and if
it detects a gc.log file, it runs `git gc --aggressive`.
There shouldn't be any reason for people to have to think about that
aspect. Remove that info from the README and make it lighter weight.
Signed-off-by: Théo Lebrun <theo.lebrun@bootlin.com>
utils/pack-repositories did the following on repos which have a gc.log
file existing (created when GC fails):
git prune
git gc --aggressive
git prune
git gc --aggressive
Here we:
- Delete utils/pack-repositories; we don't want that detection to be
done manually. Instead, we integrate the gc.log detection into
utils/index that should be called often.
- Create a hidden flag ($ELIXIR_GC) to allow manual trigger.
- Replace the above sequence with a simpler `git gc --aggressive`.
Let's trust Git.
- Do a `git gc --auto` in the default case. This call is automatically
done by porcelain commands but we don't run any so let's give Git an
opportunity to cleanup from time to time (heuristic based).
- Replace the gc.log detection from:
find . -name gc.log
To:
test -e $data/$project/repo/gc.log
It should be more reliable. With the first approach we risk projects
that contain a file gc.log to trigger the detection on each run.
Signed-off-by: Théo Lebrun <theo.lebrun@bootlin.com>
Avoid the following Git warning:
hint: Using 'master' as the name for the initial branch. This default branch name
hint: is subject to change. To configure the initial branch name to use in all
hint: of your new repositories, which will suppress this warning, call:
hint:
hint: git config --global init.defaultBranch <name>
hint:
hint: Names commonly chosen instead of 'master' are 'main', 'trunk' and
hint: 'development'. The just-created branch can be renamed via this command:
hint:
hint: git branch -m <name>
Signed-off-by: Théo Lebrun <theo.lebrun@bootlin.com>
Previously:
LXR_PROJ_DIR=/srv/elixir-data ./utils/update-elixir-data
Now:
./utils/index /srv/elixir-data --all
The impact is slightly different: it also has the side-effect of
creating all known projects (Linux, U-Boot, etc.) if they didn't exist.
We have asked around and we are not aware of any other Elixir instance.
To keep the previous behavior, if people don't want to index all
supported projects:
x=/srv/elixir-data
find $x -mindepth 1 -maxdepth 1 -printf "%f\n | \
xargs -L1 -r ./utils/index $x
Signed-off-by: Théo Lebrun <theo.lebrun@bootlin.com>
Allow calling like:
./utils/index musl
That will do the same thing as before (fetch+index).
It works only if a previous call was made to add remotes.
Signed-off-by: Théo Lebrun <theo.lebrun@bootlin.com>
$ELIXIR_THREADS fallback to nproc is straight forward code, much more
than the incantation to find the path to the Elixir install path.
Remove the incantation and replace by simple code:
if test -z "$ELIXIR_THREADS"; then
ELIXIR_THREADS="$(nproc)"
fi
Signed-off-by: Théo Lebrun <theo.lebrun@bootlin.com>
Make utils/index-repository idempotent, meaning we can call it multiple
times on the same repo and same remotes without issues.
Also allow adding new remotes to an existing repo.
Signed-off-by: Théo Lebrun <theo.lebrun@bootlin.com>
Simplify the script. We never `cd` into the directory, we instead use
`git -C`. Avoid repeating it by creating a $git variable.
Signed-off-by: Théo Lebrun <theo.lebrun@bootlin.com>
This is pretty useful as update-elixir-data gets called often to check
for new updates. Most often, there are none, so checking all remotes at
the same time is useful. This only applies to the kernel, that is the
only project using multiple (three) remotes.
Signed-off-by: Théo Lebrun <theo.lebrun@bootlin.com>
Crawl-Delay does have a big impact on the loadavg of the server,
meaning:
- (1) most requests are from crawlers and,
- (2) most crawlers listen to Crawl-Delay.
The prod server can handle the current loadavg just fine, let's let them
up their game.
Signed-off-by: Théo Lebrun <theo.lebrun@bootlin.com>
lib.getFileFamily() returns None for many files. Our assert to ensure
the family is valid should only be done once we have checked the family
is NOT None.
Signed-off-by: Théo Lebrun <theo.lebrun@bootlin.com>
Add a list (CACHED_DEFINITIONS_FAMILIES) that tells us which families
have their definitions cached. We use that to assert at DB.__init__()
and q.query('file') that everything is working as expected.
If someone modifies lib.getFileFamily() for example, we'll get an
explicit warning that we should add a definitions cache to that new
family.
Signed-off-by: Théo Lebrun <theo.lebrun@bootlin.com>
By default it never restarts. It is a good idea to avoid leaks across
weeks of a Python process running.
Also remove useless comment about processes value (16 is probably higher
than CPU count and will be fine).
Signed-off-by: Théo Lebrun <theo.lebrun@bootlin.com>
`./script.sh get-latest <offset>` gets the full list of tags, filters
it, sorts it then returns a single result. On the Python side, it gets
the first one. If that works, it uses it, else it tries the second one,
etc.
That is a weird implementation: modify get-latest to return all tags so
that Python code only has to spawn a single subprocess.
Also, rename it from `get-latest` to `get-latest-tags`. This makes
things more explicit (what latest?) and also explicits that more than
one tag is required.
Also, argument 1 is supposed to be an offset.
No custom implementation of get_latest() did implement that.
Signed-off-by: Théo Lebrun <theo.lebrun@bootlin.com>
This code was required because we messed-up in the past regarding
caching headers. This is not required anymore because the caching set
has expired, so no well behaving user-agent should have remains.
This represents something like 240k requests to the backend (not the
cache) over two weeks. Server load was minimal because generating those
responses is really fast.
Signed-off-by: Théo Lebrun <theo.lebrun@bootlin.com>
Apparently we might need this for some C code extensions:
> Forcing a WSGI application to run within the first interpreter can be
> necessary when a third party C extension module for Python has used
> the simplified threading API for manipulation of the Python GIL and
> thus will not run correctly within any additional sub interpreters
> created by Python.
https://www.modwsgi.org/en/latest/configuration-directives/WSGIApplicationGroup.html
Signed-off-by: Théo Lebrun <theo.lebrun@bootlin.com>
Lookup if a definition exists is taking too long to render source code.
Generate small databases that only tell us if a definition exists for a
given family. Because the database is much smaller, it is faster to
query.
Many URLs could only be queried at 12 req/s. With that patch, I can do
>80 req/s on the same URLs, with the same config.
We generate the caches from update.py. We also add an edge-case to
generate the files (if they don't exist) even if no new tag exists.
Signed-off-by: Théo Lebrun <theo.lebrun@bootlin.com>
This behavior isn't used at the moment, but is much more sensible that
the past that was to give None to whatever content-type lambda we had
(eg DefList that would call split on the None).
Signed-off-by: Théo Lebrun <theo.lebrun@bootlin.com>
Before, we copied all source files at the start of the Dockerfile; that
meant we always rebuilt most steps. Only copy requirements.txt first,
then copy the rest after many steps.
Signed-off-by: Théo Lebrun <theo.lebrun@bootlin.com>