Instances (/instances/)¶
Instances sit at the core of Shaken Fist's functionality, and are the component which ties most of the other concepts in the API together. Therefore, they are also the most complicated part of Shaken Fist to explain. This description is broken into basic functionality -- showing information about instances -- and then moves onto more advanced topics like creation, deletion, and other lifecycle events.
Note
For a detailed reference on the state machine for instances, see the developer documentation on object states.
Fetching information about an instance¶
There are two main ways to fetch information about instances -- you can list all instances visible to your authenticated namespace, or you can collect information about a specific instance by providing its UUID or name.
Info
Note that the amount of information visible in an instance response will change over the lifecycle of the instance -- for example when you first request the instance be created versus when the instance has had its disk specification calculated.
REST API calls
- GET /instances: List all instances visible to your authenticated namespace.
- GET /instances/{instance_ref}: Get information about a specific instance.
Python API client: get all visible instances
Python API client: get information about a specific instance
Instance creation¶
Instance creation is by far the most complicated call in Shaken Fist in terms of the arguments that it takes. The code in the Python command line client is helpful if you need a fully worked example of every possible permutation. The OpenAPI documentation at https://openapi.shakenfist.com/#/instances/post_instances provides comprehensive and up to date documentation on all the arguments to the creation call.
The instance creation API call also takes three data structures: the diskspec;
the networkspec; and the videospec. These structures are not well documented
in the OpenAPI interface, so are documented here instead.
Since v0.8 all three are described to the API's input validation layer instead of
being handed to the handler unexamined, which has two consequences for every one
of them. A key which is not listed below is refused with a 400 naming it,
where it used to be silently discarded -- typing siz instead of size got you
a default sized disk and no indication that anything had gone wrong. And a value
of the wrong type is refused with a 400 naming the key and the element it is in
(disk[0].size: Not a valid integer.), where many of them used to reach code
which could not cope and produce a server error naming nothing.
Any of the values below may also be sent as JSON null, which means "not
supplied" and gets you the same behaviour as omitting the key -- the shipped
client and the shipped ansible collection both do that for several of them on an
ordinary request. The exceptions are the three values a handler requires: a
networkspec's network_uuid, and a videospec's model and memory. For
those a null is refused exactly as an omission is, which is the same statement
as the rest of this paragraph rather than a different one -- omitting them is
already a 400.
diskspec¶
A diskspec consists of the following fields as a JSON dictionary:
- size (integer): the size of the disk in gigabytes. Must be zero or greater; a negative size is refused with a 400, because it corrupts the cluster's capacity accounting. Omit this value, or send null, for a disk the size of its base image.
- base (string): the base image for the disk. This can be a variety of URL-like strings, as documented on the artifacts page in the user guide. For a blank disk, omit this value, send null, or send the literal string "none" -- all three mean the same thing here.
- bus (enum): the hardware bus the disk device should be attached to on the instance. In general you shouldn't care about this and can omit this value. However, in some cases, such as unmodified Microsoft Windows images it is required. The options available here are: sata; scsi; usb; virtio (the default); and nvme. While ide was previously supported, that support was removed in v0.7 due to extremely poor performance, and a bus of ide is refused with a 400 along with every other value outside that list.
- type (enum): the type of device, either "disk" (the default) or "cdrom". Any other value is refused with a 400. Before v0.8 any string was accepted here and handed to libvirt as a device name, which meant that libvirt's own device types were accepted and then behaved as plain disks, because only "cdrom" is treated specially by Shaken Fist.
A diskspec must ask for something: one which specifies neither a size nor a
base -- including a size of zero with no base, and a base of the literal
string "none", which means no base -- is refused with a 400, because it
describes a disk nobody asked for.
A full example of a diskspec is therefore:
networkspec¶
Similarly, a networkspec consists of the following fields in a JSON dictionary:
- network_uuid (string, required): the network the interface should exist on.
Despite the name this is not restricted to being a UUID -- the unique name of a
network in the namespace you are creating in is accepted as well, which is why
this value is published as a plain string rather than as a UUID. A
networkspecwith nonetwork_uuid, or with a null one, is refused with a 400. - macaddress (string): the MAC address of the interface, in the colon separated form
02:00:00:ea:3a:28. Either case is accepted. A value in any other form is rejected with a 400. Omit this value, or send null, to be allocated a MAC address automatically -- which is what the shipped client does on every call. - address (string): the IPv4 address to assign to the interface, or the literal string "none" for an interface with no address at all. Omit this value, or send null, to be allocated a random address. Because of the "none" case this value is published as a plain string and is not validated as an IPv4 address.
- model (string): the model of the network interface card. In general you should not have
to set this, although it can matter in some cases, such as unmodified Microsoft
Windows images. virtio is the default and is almost always the right answer;
e1000, rtl8139, pcnet and the i825xx family are the usual choices for a guest
without virtio drivers. There is deliberately no enumeration of legal values
here, and a value this API does not recognise is not refused: the set which
actually works is whatever the hypervisor's qemu build supports, which varies by
node and by release and which this API cannot know. If a model is wrong you will
find out when the instance fails to start, not when you create it. The one
check the value does get is the published pattern
^[A-Za-z0-9._-]{1,32}$, which every model name fits: the value is rendered into the instance's libvirt domain XML, so a value carrying an XML metacharacter is refused with a 400 (and is escaped at render time regardless). - float (boolean): whether to associate a floating IP with this interface to enable external
accessibility to the instance. Note that you can float and unfloat an interface
after instance creation if desired. A JSON boolean is the expected form, and
the shipped client and ansible collection both send one. A range of string
spellings (
true/false,yes/no,on/off,1/0, and some single-letter and case variants of each) are also read with the meaning they carry, so"false"does not float the interface. A spelling outside those sets --"tRue"is one -- is refused with a 400 under the shippedenforcedefault, but read as true under thewarn/offrollback, where the handler's ownbool()fallback answers instead of the schema check. Do not rely on either; send a real JSON boolean.
The same structure is passed to POST /instances/{instance_ref}/interfaces when hot plugging an interface into a running instance, where it is a single dictionary rather than a list of them, and is checked identically.
videospec¶
A videospec differs from a diskspec and a networkspec in that it is not
passed as a list. You only have one videospec per instance. Once again, a
videospec is a JSON dictionary with the following fields:
- model (string): the model of the video card to attach to the instance. cirrus is
the default, and vga and qxl are the other usual choices -- qxl is the one to
pair with SPICE. As with a
networkspec's model there is deliberately no enumeration of legal values, and an unrecognised one is not refused, for the same reason: the value is rendered into the libvirt domain XML, so the working set belongs to the hypervisor's qemu build rather than to this API. It carries the same^[A-Za-z0-9._-]{1,32}$pattern as thenetworkspec's model, for the same reason. - memory (integer): the amount of video RAM the video card should have, in kibibytes (blocks of 1024 bytes). A value with a fractional part is refused with a 400.
- vdi (enum): the VDI protocol to use. Options are "vnc", "spice" (the default),
"spiceconcurrent", or "spicedebug". spice and spiceconcurrent are the same except
that spiceconcurrent allows limited multi-user sessions, with subsequent sessions
not experiencing full VDI functionality. spicedebug behaves like spice but also
sets
G_MESSAGES_DEBUG=allin the qemu environment so the SPICE server emits verbose logs to the qemu log, which is useful for diagnosing connectivity issues. This enumeration is Shaken Fist's own, rather than the hypervisor's, so unlike the two model fields a value outside it is refused with a 400 when you create the instance. It previously reached the instance and then failed in a console request, long after the call which accepted it.
If you supply a videospec at all it must contain both model and memory, or
the call is refused with a 400 -- and a null counts as not containing it, since a
stored null reaches the hypervisor and produces an instance which cannot start.
Omit the whole structure, send null, or send an empty dictionary to get the
defaults instead. A null vdi is defaulted like an absent one.
REST API calls
- POST /instances/: Create an instance.
Python API client: create and then delete a simple instance
Adding network interfaces after instance creation¶
As of Shaken Fist v0.8, it is also possible to add network interfaces to an existing instance, assuming that your guest operating system supports device hot plugging (all modern Linux versions do).
REST API calls
Python API client: hot plug a network interface
Note that this example assumes the instance is running an image with the Shaken Fist in guest agent installed.
```python from shakenfist_client import apiclient
sf_client = apiclient.Client()
Create a network to hot plug to¶
hotnet = sf_client.allocate_network('10.0.0.0/24', True, True, 'hotplug')
...
Hot plug the interface in¶
netdesc = { 'network_uuid': hotnet['uuid'], 'address': '10.0.0.5', 'macaddress': '02:00:00:ea:3a:28' } sf_client.add_instance_interface(inst['uuid'], netdesc)
Other instance lifecycle operations¶
A variety of other lifecycle operations are available on instances, including deletion, and power management.
The power management actions available are:
- soft reboot: gracefully request a reboot from the instance operating system via ACPI. This is not guaranteed to actually work, but if it does is much less likely to cause disk corruption on the instance.
- hard reboot: the equivalent of holding the reset switch down on a physical machine until it reboots without operating system involvement.
- power on: turn the instance on, as if the power switch was pressed. Since v0.8 power on operations have the side effect of creating the config drive if one is specified by the instance configuration. That is, you can recreate the config drive by powering the instance off and then on again.
- power off: turn the instance immediately off, as if the power switch was held down on a physical machine.
- pause: suspend execution of the instance, but leave it hot in RAM ready to restart.
- unpause: unsuspend execution of the instance.
REST API calls
- DELETE /instances/{instance_ref}: Delete an instance.
- DELETE /instances/: Delete all instances in a namespace.
- POST /instances/{instance_ref}/rebootsoft: Soft (ACPI) reboot the instance.
- POST /instances/{instance_ref}/reboothard: Hard (reset switch) reboot the instance.
- POST /instances/{instance_ref}/poweron: Power the instance on.
- POST /instances/{instance_ref}/poweroff: Power the instance off, as if holding the power switch down.
- POST /instances/{instance_ref}/pause: Pause an instance.
- POST /instances/{instance_ref}/unpause: Unpause an instance.
Python API client: create and then delete a simple instance
Python API client: attempt a soft reboot, and hard reboot if required
Note that this example assumes the instance is running an image with the Shaken Fist in guest agent installed.
import time
from shakenfist_client import apiclient
import sys
sf_client = apiclient.Client()
i = sf_client.create_instance(
'example', 1, 1024, None,
[{
'size': 20,
'base': 'debian:11',
'bus': None,
'type': 'disk'
}],
None, None, side_channels=['sf-agent'])
# Wait for the instance to be created, or error out. Use instance events to
# provide status updates during boot.
while i['state'] not in ['created', 'error']:
events = sf_client.get_instance_events(i['uuid'])
print('Waiting for the instance to start: %s' % events[0]['message'])
time.sleep(5)
i = sf_client.get_instance(i['uuid'])
# Check the instance is created correctly
if i['state'] != 'created':
print('Instance is not in a created state!')
sys.exit(1)
print('Instance is created')
# Wait for the agent to report the reboot time
while not i['agent_system_boot_time']:
print('Waiting for agent to start: %s' % i['agent_state'])
time.sleep(20)
i = sf_client.get_instance(i['uuid'])
initial_boot = i['agent_system_boot_time']
print('Instance booted at %d' % initial_boot)
# Now try to soft reboot the instance, wait up to 60 seconds for a reboot to
# be detected
sf_client.reboot_instance(i['uuid'], hard=False)
print('Soft rebooting instance')
time.sleep(60)
i = sf_client.get_instance(i['uuid'])
# Wait for the agent to report the reboot time again
while not i['agent_system_boot_time']:
print('Waiting for agent to start: %s' % i['agent_state'])
time.sleep(20)
i = sf_client.get_instance(i['uuid'])
if i['agent_system_boot_time'] != initial_boot:
print('Boot time changed from %d to %s'
% (initial_boot, i['agent_system_boot_time']))
else:
# We failed to soft reboot, let's hard reboot instead
sf_client.reboot_instance(i['uuid'], hard=True)
print('Instance did not reboot, hard rebooting')
Sample output:
$ python3 example.py
Waiting for the instance to start: schedule complete
Instance is created
Waiting for agent to start: not ready (no contact)
Waiting for agent to start: not ready (no contact)
Waiting for agent to start: not ready (no contact)
Instance booted at 1684404969
Soft rebooting instance
Boot time changed from 1684404969 to 1684405036.0
Python API client: power off and then on an instance
import time
from shakenfist_client import apiclient
import sys
sf_client = apiclient.Client()
i = sf_client.create_instance(
'example', 1, 1024, None,
[{
'size': 20,
'base': 'debian:11',
'bus': None,
'type': 'disk'
}],
None, None)
# Wait for the instance to be created, or error out. Use instance events to
# provide status updates during boot.
while i['state'] not in ['created', 'error']:
events = sf_client.get_instance_events(i['uuid'])
print('Waiting for the instance to start: %s' % events[0]['message'])
time.sleep(5)
i = sf_client.get_instance(i['uuid'])
# Check the instance is created correctly
if i['state'] != 'created':
print('Instance is not in a created state!')
sys.exit(1)
print('Instance is created')
# Check the instance is created correctly
if i['power_state'] != 'on':
print('Instance is not in powered on state!')
sys.exit(1)
# Power the instance off
sf_client.power_off_instance(i['uuid'])
while i['power_state'] != 'off':
print('Waiting for the instance to power off')
time.sleep(5)
i = sf_client.get_instance(i['uuid'])
time.sleep(30)
# Power the instance on
sf_client.power_on_instance(i['uuid'])
while i['power_state'] != 'on':
print('Waiting for the instance to power on')
time.sleep(5)
i = sf_client.get_instance(i['uuid'])
print('Done')
Python API client: pause and unpause an instance
Other instance information¶
We can also request other information for an instance. For example, we can list the instance's network interfaces, or the events for the instance. See the user guide for a general introduction to the Shaken Fist event system.
REST API calls
- GET /instances/{instance_ref}/interfaces: Request information on the instance's network interfaces, if any.
- GET /instances/{instance_ref}/events: Fetch events for a specific instance.
Python API client: list network interfaces for an instance
Note that the interface details for an instance wont be populated until the instance has started being created on the hypervisor node. Specifically, this can be some time later if an image needs to be fetched from the Internet and transcoded. Therefore in this example we wait for the instance to be created before displaying interface details.
import json
from shakenfist_client import apiclient
import time
sf_client = apiclient.Client()
i = sf_client.create_instance(
'example', 1, 1024, None,
[{
'size': 20,
'base': 'debian:11',
'bus': None,
'type': 'disk'
}],
None, None)
# Wait for the instance to be created, or error out. Use instance events to
# provide status updates during boot.
while i['state'] not in ['created', 'error']:
events = sf_client.get_instance_events(i['uuid'])
print('Waiting for the instance to start: %s' % events[0]['message'])
time.sleep(5)
i = sf_client.get_instance(i['uuid'])
# Check the instance is created correctly
if i['state'] != 'created':
print('Instance is not in a created state!')
sys.exit(1)
print('Instance is created')
# Fetch and display interface details
ifaces = sf_client.get_instance_interfaces(i['uuid'])[0]
print(json.dumps(ifaces, indent=4, sort_keys=True))
$ python3 example.py
Waiting for the instance to start: Fetching required blob ffdfce7f-728e-4b76-83c2-304e252f98b1, 30% complete
Instance is created
[
{
"floating": null,
"instance_uuid": "d512e9f5-98d6-4c36-8520-33b6fc6de15f",
"ipv4": "10.0.0.6",
"macaddr": "02:00:00:73:18:66",
"metadata": {},
"model": "virtio",
"network_uuid": "6aaaf243-0406-41a1-aa13-5d79a0b8672d",
"order": 0,
"state": "created",
"uuid": "b1981e81-b37a-4176-ba37-b61bc7208012",
"version": 3
}
]
Python API client: list events for an instance
import json
from shakenfist_client import apiclient
sf_client = apiclient.Client()
interfaces = sf_client.get_instance_events('c0d52a77-0f8a-4f19-bec7-0c05efb03cb4')
print(json.dumps(interfaces, indent=4, sort_keys=True))
Note that events are returned in reverse chronological order and are limited to the 100 most recent events.
[
...
{
"duration": null,
"extra": {
"cpu usage": {
"cpu time ns": 357485828000,
"system time ns": 66297716000,
"user time ns": 291188112000
},
"disk usage": {
"vda": {
"actual bytes on disk": 956301312,
"errors": -1,
"read bytes": 406776320,
"read requests": 12225,
"write bytes": 2105954304,
"write requests": 3657
},
"vdb": {
"actual bytes on disk": 102400,
"errors": -1,
"read bytes": 279552,
"read requests": 74,
"write bytes": 0,
"write requests": 0
}
},
"network usage": {
"02:00:00:1d:24:ae": {
"read bytes": 147084732,
"read drops": 0,
"read errors": 0,
"read packets": 16484,
"write bytes": 2166754,
"write drops": 0,
"write errors": 0,
"write packets": 13144
}
}
},
"fqdn": "sf-2",
"message": "usage",
"timestamp": 1685229509.9592097,
"type": "usage"
},
...
]
Out-of-band interactions with instances¶
Shaken Fist supports three types of instance consoles, which provide out-of-band management of instances -- that is, the instance does not need to have functioning networking for these consoles to work. You can read a general introduction of Shaken Fist's console functionality in the user guide. This page focuses on the API calls which are used to implement the console functionality in the Shaken Fist client.
- Read only console: to download the most recent portion of the read only text
serial console, or clear the console, use the
/instances/{instance_ref}/consoledataAPI calls below. - Interactive serial console: lookup the console port from the instance details fetch (as described above), and then connect to that port on the hypervisor node with a TCP client such as telnet.
- Interactive VDI console: lookup the VDI console port from the instance details
fetch (as described above), and then connect to that port on the hypervisor
with the correct client (currently one of VNC or SPICE). Alternatively, use
the
/instances/{instance_ref}/vdiconsolehelperAPI call described below to download avirt-viewerconfiguration file and then connect withvirt-viewer. See the example below for more details. - Proxied VDI console: since v0.8, if the cluster operator has enabled the
Kerbside integration, use the
/instances/{instance_ref}/vdiconsoleproxyAPI call below to mint a short lived signed token and receive a proxy URL of the form<KERBSIDE_URL>/sf-console.vv?token=<jwt>. The user's viewer redeems this URL against the Kerbside proxy, which validates the token offline and relays the SPICE session, so no direct network access to the hypervisor is needed. See the VDI console tokens operator guide for how the integration is enabled and operated.
REST API calls
- GET /instances/{instance_ref}/consoledata: Fetch read only serial console data for an instance
- DELETE /instances/{instance_ref}/consoledata: Clear the read only serial console for an instance.
- GET /instances/{instance_ref}/vdiconsolehelper: Generate and return a
virt-viewerconfiguration file for connecting to the interactive VDI console for the instance (if configured). - GET /instances/{instance_ref}/vdiconsoleproxy: Mint a short lived Kerbside VDI console token and return a proxy URL (
{url, expires_at}) for the SPICE console of the instance. Returns 404 if the Kerbside integration is not configured, 406 if the instance is not in thecreatedstate, 409 if the instance does not have a SPICE console, and 500 if no signing key is configured.
Python API client: connect seamlessly to a VDI console using virt-viewer
import os
from shakenfist_client import apiclient
import subprocess
import tempfile
import time
sf_client = apiclient.Client()
i = sf_client.create_instance(
'example', 1, 1024, None,
[{
'size': 20,
'base': 'debian:11',
'bus': None,
'type': 'disk'
}],
None, None)
# Wait for the instance to be created, or error out. Use instance events to
# provide status updates during boot.
while i['state'] not in ['created', 'error']:
events = sf_client.get_instance_events(i['uuid'])
print('Waiting for the instance to start: %s' % events[0]['message'])
time.sleep(5)
i = sf_client.get_instance(i['uuid'])
# Check the instance is created correctly
if i['state'] != 'created':
print('Instance is not in a created state!')
sys.exit(1)
print('Instance is created')
# We don't use NamedTemporaryFile as a context manager as the .vv file
# will also attempt to clean up the file.
(temp_handle, temp_name) = tempfile.mkstemp()
os.close(temp_handle)
try:
with open(temp_name, 'w') as f:
f.write(sf_client.get_vdi_console_helper(i['uuid']))
p = subprocess.run('remote-viewer %s' % temp_name, shell=True)
print('Remote viewer process exited with %d return code' % p.returncode)
finally:
if os.path.exists(temp_name):
os.unlink(temp_name)
Executing commands within an instance¶
Since v0.7, assuming a given instance has the Shaken Fist agent installed and
running, and was created with a sf-agent side channel, you can use the Shaken
Fist agent to move data into and out of the instance and execute commands without
the instance needing to have working networking configured. You can read more
about the exact requirements for agent connectivity in the
API reference guide for agent operations.
Agent Operations are not created directly -- they are a side effect of a call to one of the API methods below, which create an Agent Operation so the caller can track the state of their request. At the time of writing, you can perform the following operations via the agent:
- copy the contents of a blob into an instance and change its file permissions. The python API client has a helper to upload the file into a blob before copying to the instance.
- execute a command and return its results (exit code, stdout, stderr).
- get the contents of a file within an instance into a blob.
REST API calls
- POST /instances/{instance_ref}/agent/execute: execute a command within an instance and return results.
- POST /instances/{instance_ref}/agent/put: copy a blob into an instance at the specified location with the specified permissions.
Bounding how long an agent operation may take¶
All three creating calls accept an optional deadline_seconds, and
agent/get and agent/put additionally accept an optional
progress_timeout_seconds. Both are counts of seconds, and both refuse
a negative value with a 400. Both are also capped by the operator
ceiling AGENT_OPERATION_MAX_DEADLINE (86400 seconds, one day, unless
the operator has changed it), published as the parameters' maximum in
the API specification and likewise refused with a 400 above it.
deadline_seconds is how long the operation may continue to be
dispatched or execute, counted from the moment the API server received
your request rather than from when the agent picked the work up. Time
spent queued behind another operation on the same instance, and any
preflight work such as fetching a blob onto the hypervisor, both count
against it.
progress_timeout_seconds is how long the operation may go without
making forward progress. It applies to commands which can report
progress -- the transfers behind agent/get and agent/put -- and is
the more useful of the two for a large file, because a transfer can be
perfectly healthy and still take a long time.
Each parameter has three meanings, and the difference between the first two matters:
| You send | What happens |
|---|---|
| nothing | The server default applies: AGENT_OPERATION_DEFAULT_DEADLINE (600 seconds) or AGENT_OPERATION_DEFAULT_PROGRESS_TIMEOUT (30 seconds), unless your operator has changed them. |
0 |
None at all. The operation is not bounded by that mechanism. |
| a positive number | That many seconds, up to AGENT_OPERATION_MAX_DEADLINE. |
Sending 0 is not the same as omitting the parameter. Streaming a very
large file out of an instance is the case the distinction exists for:
send deadline_seconds: 0 with a progress_timeout_seconds, and the
transfer is allowed to take as long as it takes while a genuine stall
is still detected.
There is one exception to 0 meaning unbounded: an operation with
both budgets disabled -- deadline_seconds: 0 on agent/execute,
which always stores no progress timeout, or combined with an explicit
progress_timeout_seconds: 0 on the transfers -- would otherwise hold
its instance's single executor slot for as long as the guest command
cared to run, blocking every other agent operation against that
instance. Such an operation is expired AGENT_OPERATION_MAX_DEADLINE
seconds after it last changed state instead.
agent/execute does not accept progress_timeout_seconds, and sending
it is a 400. Nothing an executed command does is observable as
progress, so a timeout there could never fire; only deadline_seconds
is meaningful.
What happens when a budget runs out
An operation which exhausts either budget, and for which no retry is
possible (see below), moves to expired, a terminal state distinct
from error. error means the operation itself failed; expired
means a budget you set ran out, and the operation's expiry_reason
field says which: deadline or progress. A deadline expiry is
answered by asking for a longer deadline; a progress expiry means
the agent stalled, which a bigger deadline does not fix.
A wall-clock deadline running out is always final -- there is no
time left for a further attempt to deliver anything in -- but a
progress-timeout stall may be retried a bounded number of times
before landing in expired; see
Agent Operations
in the operator guide for the retry rules and the node-local reaper
that backs them up.
Three places enforce this, so an abandoned operation is retired
wherever it happens to be sitting: when it reaches the head of the
instance's queue, during preflight (either side of any blob copy),
and once per second while the executor is running it. Only the
executor can enforce progress_timeout_seconds, since it is the
only one of the three watching replies arrive.
The reason is recorded as an audit event against both the operation and its instance, and the instance's copy is the one which survives: an expired operation is eventually hard deleted, like a completed one.
Python API client: execute a command on an instance
import json
from shakenfist_client import apiclient
sf_client = apiclient.Client()
agentop = sf_client.instance_execute('...uuid...', 'cat /etc/os-release')
print(json.dumps(agentop, indent=4, sort_keys=True))
Which would return something along the lines of:
{
"attempts": 0,
"commands": [
{
"block-for-result": true,
"command": "execute",
"commandline": "cat /etc/os-release"
}
],
"deadline": 1787428090.5,
"expiry_reason": null,
"instance_uuid": "a771fb13-aaad-4cb6-a86b-7ee51e7bacc6",
"last_progress": null,
"metadata": {},
"namespace": "vdi",
"progress_timeout": 0.0,
"results": {
"0": {
"command-line": "cat /etc/os-release",
"result": true,
"return-code": 0,
"stderr": "",
"stdout": "PRETTY_NAME=\"Debian GNU/Linux 11 (bullseye)\"..."
}
},
"state": "complete",
"uuid": "93fb538c-84f5-4ff8-83ba-2be5f5f92954",
"version": 3
}
Python API client: put a file onto an instance via a blob
import os
from shakenfist_client import apiclient
from shakenfist_client import util
sf_client = apiclient.Client()
if not sf_client.check_capability('blob-search-by-hash'):
blob = None
else:
# We can cheat here -- if we already have a blob in the cluster with the
# checksum of the file we're uploading, we can skip the upload entirely and
# just reuse that blob.
blob = util.checksum_with_progress(sf_client, 'README.md')
if not blob:
artifact = util.upload_artifact_with_progress(
sf_client, 'upload-to-instance', 'README.md', None)
else:
print('Recycling existing blob')
artifact = sf_client.blob_artifact(
'upload-to-instnace', blob['uuid'], source_url=None)
print('Created artifact %s' % artifact['uuid'])
st = os.stat('README.md')
sf_client.instance_put_blob(
'...instance_ref...', artifact['blob_uuid'], '/tmp/README.md', st.st_mode)
Which would return something along the lines of:
Python API client: get a file from an instance via a blob
from shakenfist_client import apiclient
import sys
sf_client = apiclient.Client()
op = sf_client.instance_get('...instance_ref...', '/tmp/README.md')
if '0' not in op.get('results', {}):
print('Results not available.')
sys.exit(1)
blob_uuid = op['results']['0'].get('content_blob')
if not blob_uuid:
print('Results did not include content')
sys.exit(1)
with open('/tmp/README.md', 'wb') as f:
for chunk in sf_client.get_blob_data(blob_uuid):
f.write(chunk)
Fetching information about an Instance's Agent Operations¶
Additionally, you can list the agent operations for a given instance.
REST API calls
- GET /instances/{instance_ref}/agentoperations: List all agent operations for an instance.
Python API client: get all agent operations for a specific instance
The instance here had the following command line commands run before this sample script was run:
$ sf-client instance upload ...uuid... README.md /tmp/README.md
$ sf-client --simple instance execute ...uuid... "cat /tmp/README.md"
import json
from shakenfist_client import apiclient
sf_client = apiclient.Client()
agentops = sf_client.get_instance_agentoperations('...uuid...', all=True)
print(json.dumps(agentops, indent=4, sort_keys=True))
Note the all argument here. By default you are only returned agent operations
which are queued to execute. To see all agent operations including those
which have completed execution, pass all=True. This script outputs:
[
{
"attempts": 0,
"commands": [
{
"blob_uuid": "09306f15-b1b3-4850-afb4-f4179559fa7f",
"command": "put-blob",
"path": "/tmp/README.md"
},
{
"command": "chmod",
"mode": 33188,
"path": "/tmp/README.md"
}
],
"deadline": 1787428090.5,
"expiry_reason": null,
"instance_uuid": "a771fb13-aaad-4cb6-a86b-7ee51e7bacc6",
"last_progress": null,
"metadata": {},
"namespace": "vdi",
"progress_timeout": 30.0,
"results": {
"0": {
"path": "/tmp/README.md"
},
"1": {
"path": "/tmp/README.md"
}
},
"state": "complete",
"uuid": "343049d7-da2a-46f2-bb5c-edb783ec1fb9",
"version": 3
},
{
"attempts": 0,
"commands": [
{
"block-for-result": true,
"command": "execute",
"commandline": "cat /tmp/README.md"
}
],
"deadline": 1787428090.5,
"expiry_reason": null,
"instance_uuid": "a771fb13-aaad-4cb6-a86b-7ee51e7bacc6",
"last_progress": null,
"metadata": {},
"namespace": "vdi",
"progress_timeout": 0.0,
"results": {
"0": {
"command-line": "cat /tmp/README.md",
"result": true,
"return-code": 0,
"stderr": "",
"stdout": "...content of file..."
}
},
"state": "complete",
"uuid": "5a00d6f3-19b6-42bc-b1df-ddc4e5a299e9",
"version": 3
}
]
Console screen captures¶
Since v0.8, Shaken Fist has provided an API for collecting screen captures of the instance console. This works for either serial consoles or graphical consoles, its literally the same was whatever would have been displayed on the monitor if the instance was a physical machine.
REST API calls
- GET /instances/{instance_ref}/screenshot: Collect a screenshot for an instance.
This API call returns a blob UUID, you then need to collect the contents of the blob using the GET /blobs/{blob_uuid}/data API call. The python Shaken Fist API client perfoms both operations for you and returns an iterator of binary chunks ready for you to process or write to a file.
Python API client: collect a screenshot for an instance an instance
Object References¶
Instance API responses include references_to and references_from fields that
show the relationships between instances and other objects in the system. These
fields help you understand how instances are connected to blobs and other objects.
The references_to field shows what objects reference this instance (typically
empty for instances). The references_from field shows what blobs this instance
references (e.g., disk blobs, NVRAM template blobs).
Example references_from output for an instance
"references_from": {
"disk": [
{
"source_object_type": "instance",
"source_uuid": "d51aa352-368c-484c-9e4c-4542927b4277",
"relationship": "disk",
"relationship_value": "0",
"target_object_type": "blob",
"target_uuid": "5117f778-b214-4184-8358-f2c7376b76db",
"created": 1683995934.357137,
"last_active": 1684054381.217045
}
],
"nvram_template": [
{
"source_object_type": "instance",
"source_uuid": "d51aa352-368c-484c-9e4c-4542927b4277",
"relationship": "nvram_template",
"relationship_value": null,
"target_object_type": "blob",
"target_uuid": "abc123-def456-ghi789",
"created": 1683995934.357137,
"last_active": 1684054381.217045
}
]
}
Metadata¶
All objects exposed by the REST API may have metadata associated with them. This metadata is for storing values that are of interest to the owner of the resources, not Shaken Fist. Shaken Fist does not attempt to interpret these values at all, with the exception of the instance affinity metadata values. The metadata store is in the form of a key value store, and a general introduction is available in the user guide.
Info
Note that for affinity metadata to be processed by the scheduler, it must be present in the instance create API call, which is why that call takes a metadata argument. Adding affinity metadata after instance creation will not affect the placement of that instance, but would affect the placement of future instances.
REST API calls
- GET /instances/{instance_ref}/metadata: Get metadata for an instance.
- POST /instances/{instance_ref}/metadata: Create a new metadata key for an instance.
- DELETE /instances/{instance_ref}/metadata/{key}: Delete a specific metadata key for an instance.
- PUT /instances/{instance_ref}/metadata/{key}: Update an existing metadata key for an instance.
Python API client: set metadata on an instance
Python API client: get metadata for an instance
Python API client: delete metadata for an instance
Cluster operations¶
Since v0.8, cluster operations for a given instance have been exposed. This lists the various pieces of queued work that Shaken Fist has executed on a given object, with the limitation that completed cluster operations are hard deleted after CLEANER_DELAY seconds (which defaults to one hour).
REST API calls
- GET /instances/{instance_ref}/clusteroperations: Get cluster operations for an instance.