Monday, 22 August 2011

Concurrency management


1. semaphores and mutex;
context: kernel threads. the thread should be in sleep state for sometimes so as to perform its necessary tasks.
init_MUTEX
down_interruptible
up

2.completion : Allowing one thread to tell another that the job is done
context : interrupt handlers. we are doing some actions and we should expect its reaction from the hardware and which is an immediate reaction then we can use completion. example waiting for command completion interrupt from hardware.
init_completion
wait_for_completion
complete
3.spinlock
context: atomic, interrupt handlers
spin_lock_init
spin_lock_irqsave
spin_unlock_restore

spinlock : single holder lock.if you can't get the spinlock, you keep trying (spinning) until you can.The code written within the locking should be atomic (it should not sleep otherwise it should not given up the processor). If so the deadlock will occur.

semaphores: more than one holder but commonly used as single holder  lock(mutex).If you can't get a semaphore, your task will put itself on the queue, and be woken up when the semaphore is released. This means the CPU will do something else while you are waiting.

Linux Boot Process


Boot sequence summary

    * BIOS
    * Master Boot Record (MBR)
    * LILO or GRUB
    * Kernel
    * init
    * Run Levels

BIOS

Load boot sector from one of:

    * Floppy
    * CDROM
    * Hard drive

The boot order can be changed from within the BIOS. BIOS setup can be entered by pressing a key during bootup. The exact key depends varies, but is often one of Del, F1, F2, or F10.
(DOS) Master Boot Record (MBR)
DOS in the context includes MS-DOS, Win95, and Win98.

    * BIOS loads and execute the first 512 bytes off the disk (/dev/hda)
    * Standard DOS MBR will:
          o look for a primary partition (/dev/hda1-4) marked bootable
          o load and execute first 512 bytes of this partition
    * can be restored with fdisk /mbr from DOS

LILO

    * does not understand filesystems
    * code and kernel image to be loaded is stored as raw disk offsets
    * uses the BIOS routines to load

Loading sequence

    * load menu code, typically /boot/boot.b
    * prompt for (or timeout to default) partition or kernel
    * for "image=" (ie Linux) option load kernel image
    * for "other=" (ie DOS) option load first 512 bytes of the partition

Reconfiguring LILO

One minute guide to installing a new kernel

    * copy kernel image (bzImage) and modules to /boot and /lib/modules
    * edit /etc/lilo.conf
          o duplicate image= section, eg:


                  image=/bzImage-2.4.14

                    label=14

                    read-only

          o man lilo.conf for details
    * run /sbin/lilo
    * reboot to test

GRUB

    * Understands file systems
    * config lives in /boot/grub/menu.lst or /boot/boot/menu.lst

Kernel

    * initialise devices
    * (optionally loads initrd, see below)
    * mounts root filesystem
          o specified by lilo or loadin with root= parameter
          o kernel prints: VFS: Mounted root (ext2 filesystem) readonly.
    * runs /sbin/init which is process number 1 (PID=1)
          o init prints: INIT: version 2.76 booting
          o can be changed with boot= parameter to lilo, eg boot=/bin/sh can be useful to rescue a system which is having trouble booting.

initrd

Allows setup to be performed before root FS is mounted

    * lilo or loadlin loads ram disk image
    * kernel runs /linuxrc
          o load modules
          o initialise devices
          o /linuxrc exits
    * "real" root is mounted
    * kernel runs /sbin/init

Details in /usr/src/linux/Documentation/initrd.txt (part of the kernel source).
/sbin/init

    * reads /etc/inittab (see man inittab which specifies the scripts below
          o Run boot scripts:
                + debian: run /etc/init.d/rcS which runs:
                      # /etc/rcS.d/S* scripts
                      # /etc/rc.boot/* (depreciated)
                + redhat: /etc/rc.d/rc.sysinit script which: loads modules, check root FS and mount RW, mount local FS, setup network, and mount remote FS
          o switches to default runlevel eg 3.
                + run scripts /etc/rc3.d/S*
                + run programs specified in /etc/inittab

Run Levels

    * 0 halt
    * 1 single user
    * 2-4 user defined
    * 5 X11 only (0 or 1 text console)
    * 6 Reboot
    * Default is defined in /etc/inittab, eg:
          o id:3:initdefault:
    * The current runlevel can be changed by running /sbin/telinit # where # is the new runlevel, eg typing telinit 6 will reboot.

Run Level programs

    * Scripts in /etc/rc*.d/* are symlinks to /etc/init.d
          o Scripts prefixed with S will be started when the runlevel is entered, eg /etc/rc5.d/S99xdm
          o Scripts prefixed with K will be killed when the runlevel is entered, eg /etc/rc6.d/K20apache
          o X11 login screen is typically started by one of S99xdm, S99kdm, or S99gdm.
    * Run programs for specified run level
    * /etc/inittab lines:
          o 1:2345:respawn:/sbin/getty 9600 tty1
                + Always running in runlevels 2, 3, 4, or 5
                + Displays login on console (tty1)
          o 2:234:respawn:/sbin/getty 9600 tty2
                + Always running in runlevels 2, 3, or 4
                + Displays login on console (tty2)
          o l3:3:wait:/etc/init.d/rc 3
                + Run once when switching to runlevel 3.
                + Uses scripts stored in /etc/rc3.d/
          o ca:12345:ctrlaltdel:/sbin/shutdown -t1 -a -r now
                + Run when control-alt-delete is pressed

Boot Summary

    * lilo
          o /etc/lilo.conf
    * debian runs
          o /etc/rcS.d/S* scripts
          o /etc/rc3.d/S* scripts
    * redhat runs
          o /etc/rc.d/rc.sysinit script
          o /etc/rc.d/rc3.d/S* scripts

PCI Driver Flow

1. Fill up the pci_device_id table as follows.

static const struct pci_device_id pci_ids[] = {
{
.vendor = VENDOR_ID,
.device = DEVICE_ID,
.subvendor = PCI_ANY_ID,
.subdevice = PCI_ANY_ID,
.driver_data = (unsigned long) ~HOST_FORCE_PCI,
},

{   /* all zero 's */ },
};

2. This pci_device_id structure needs to be exported to user space to allow the hotplug
and module loading systems know what module works with what hardware devices.
The macro MODULE_DEVICE_TABLE accomplishes this.

An example is:
MODULE_DEVICE_TABLE(pci, pci_ids);

3. Fill the standard PCI_Driver structure with the PCI_id table,probe and remove functions.

say for example:

static struct pci_driver mypci_driver = {
.name       = DRIVER_NAME,
.id_table    = pci_ids,
.probe       = mypci_probe,
.remove     = __devexit_p(mypci_remove),
.suspend    = NULL,
.resume     = NULL,
};

 The main structure that all PCI drivers must create in order to be registered with the kernel properly is the struct pci_driver structure. This structure consists of a number of function callbacks and variables that describe the PCI driver to the PCI core.

4. Registering a PCI Driver:

**To register the struct pci_driver with the PCI core, a call to pci_register_driver (for network register_netdev,for char misc_register,for block drivers register_blkdev) is made with a pointer to the struct pci_driver.

**This is traditionally done in the module initialization code for the PCI driver:

static int pci_init(void)
{
PRINT(" \n\nPCI DRIVER IS LOADED IN KERNEL SPACE \n\n");
return pci_register_driver(&mypci_driver);
}

**Note that the pci_register_driver function either returns a negative error number
or 0 if everything was registered successfully.

5.Accessing PCI configuration space:

The configuration space can be accessed through 8-bit, 16-bit, or 32-bit data transfers at any time.
The relevant functions are prototyped in <linux/pci.h>:

int pci_read_config_byte(struct pci_dev *dev, int where, u8 *val);
int pci_read_config_word(struct pci_dev *dev, int where, u16 *val);
int pci_read_config_dword(struct pci_dev *dev, int where, u32 *val);
int pci_write_config_byte(struct pci_dev *dev, int where, u8 val);
int pci_write_config_word(struct pci_dev *dev, int where, u16 val);
int pci_write_config_dword(struct pci_dev *dev, int where, u32 val);

6. Enabling the PCI Device:

**In the probe function for the PCI driver, before the driver can access any device resource
(I/O region or interrupt) of the PCI device, the driver must call the pci_enable_device
function:

**int pci_enable_device(struct pci_dev *dev);

**This function actually enables the device. It wakes up the device and in some cases also assigns its interrupt line and I/O regions.

struct pci_dev {
        struct list_head global_list;   /* node in list of all PCI devices */
        struct list_head bus_list;      /* node in per-bus list */
        struct pci_bus  *bus;           /* bus this device is on */
        struct pci_bus  *subordinate;   /* bus this device bridges to */

        void            *sysdata;       /* hook for sys-specific extension */
        struct proc_dir_entry *procent; /* device entry in /proc/bus/pci */

        unsigned int    devfn;          /* encoded device & function index */
        unsigned short  vendor;
        unsigned short  device;
        unsigned short  subsystem_vendor;
        unsigned short  subsystem_device;
        unsigned int    class;          /* 3 bytes: (base,sub,prog-if) */
        u8              hdr_type;       /* PCI header type (`multi' flag masked out) */
        u8              rom_base_reg;   /* which config register controls the ROM */
        u8              pin;            /* which interrupt pin this device uses */

        struct pci_driver *driver;      /* which driver has allocated this device */
        u64             dma_mask;       /* Mask of the bits of bus address this
                                           device implements.  Normally this is
                                           0xffffffff.  You only need to change
                                           this if your device has broken DMA
                                           or supports 64-bit transfers.  */

        pci_power_t     current_state;  /* Current operating state. In ACPI-speak,
                                           this is D0-D3, D0 being fully functional,
                                           and D3 being off. */

        pci_channel_state_t error_state;        /* current connectivity state */
        struct  device  dev;            /* Generic device interface */

        /* device is compatible with these IDs */
        unsigned short vendor_compatible[DEVICE_COUNT_COMPATIBLE];
        unsigned short device_compatible[DEVICE_COUNT_COMPATIBLE];

        int             cfg_size;       /* Size of configuration space */

        /*
         * Instead of touching interrupt line and base address registers
         * directly, use the values stored here. They might be different!
         */
        unsigned int    irq;
        struct resource resource[DEVICE_COUNT_RESOURCE]; /* I/O and memory regions + expansion ROMs */

        /* These fields are used by common fixups */
        unsigned int    transparent:1;  /* Transparent PCI bridge */
        unsigned int    multifunction:1;/* Part of multi-function device */
        /* keep track of device state */
        unsigned int    is_enabled:1;   /* pci_enable_device has been called */
        unsigned int    is_busmaster:1; /* device is busmaster */
        unsigned int    no_msi:1;       /* device may not use msi */
        unsigned int    block_ucfg_access:1;    /* userspace config space access is blocked */

        u32             saved_config_space[16]; /* config space saved at suspend time */
        struct hlist_head saved_cap_space;
        struct bin_attribute *rom_attr; /* attribute descriptor for sysfs ROM entry */
        int rom_attr_enabled;           /* has display of the rom attribute been enabled? */
        struct bin_attribute *res_attr[DEVICE_COUNT_RESOURCE]; /* sysfs file for resources */
};


7.Enabling the PCI-Bus mastering for the device by calling pci_set_master Kernel API.
  This API enables bus-mastering on the device and calls pcibios_set_master  to do the needed arch specific settings
  It is done in the driver,if it is needed.

**BusMastering:
bus mastering is the capability of devices on the PCI bus (other than the system chipset, of course) to
  take control of the bus and perform transfers directly.PCI's design allows bus mastering of multiple devices on the bus simultaneously,
  with the arbitration circuitry working to ensure that no device on the bus (including the processor!) locks out any other device.
  At the same time though, it allows any given device to use the full bus throughput if no other device needs to transfer anything.

**PCI transactions work in a master-slave relationship. A master is an agent that initiates a transaction (can be a read or a write).
While the host CPU is often the bus master, all PCI boards can potentially claim the bus and become a bus master.

**When a PCI device is enabled, it's bus mastering is also enabled.  This occurs before any driver code is
executed.

8. Accessing Memory Regions

**unsigned long pci_resource_start(struct pci_dev *dev, int bar);
The function returns the first address (memory address or I/O port number)
associated with one of the six PCI I/O regions. The region is selected by the integer
bar (the base address register), ranging from 0–5 (inclusive).

**unsigned long pci_resource_end(struct pci_dev *dev, int bar);
The function returns the last address that is part of the I/O region number bar.
Note that this is the last usable address, not the first address after the region.

**unsigned long pci_resource_flags(struct pci_dev *dev, int bar);
This function returns the flags associated with this resource.
All resource flags are defined in <linux/ioport.h>; the most important are:
IORESOURCE_IO
IORESOURCE_MEM ( if the resource type is memory resource, then we can access the register by means of readl/b/w/writel/b/w functions)

#define pci_resource_start(dev,bar)   ((dev)->resource[(bar)].start)
#define pci_resource_end(dev,bar)     ((dev)->resource[(bar)].end)
#define pci_resource_flags(dev,bar)   ((dev)->resource[(bar)].flags)
#define pci_resource_len(dev,bar) \
       ((pci_resource_start((dev),(bar)) == 0 &&       \
         pci_resource_end((dev),(bar)) ==              \
         pci_resource_start((dev),(bar))) ? 0 :        \
                                                       \
        (pci_resource_end((dev),(bar)) -               \
         pci_resource_start((dev),(bar)) + 1))

9.Reserving the PCI I/O and memory regions.

int pci_request_region (struct pci_dev * pdev, int bar, char * res_name);

Arguments:
pdev
    PCI device whose resources are to be reserved
bar
    BAR to be reserved
res_name
    Name to be associated with resource.

Description
Mark the PCI region associated with PCI device pdev BR bar as being reserved by owner res_name.
Do not access any address inside the PCI regions unless this call returns successfully.

10. Converting physical address to virtual address.

pci.pVirtualBaseAddr=(unsigned int *)ioremap_nocache(pci.PciBaseStart,pci.length);

11.

else
    KERNELDIR ?= /lib/modules/$(shell uname -r)/build
    PWD  := $(shell pwd)
default:
    $(MAKE) -C $(KERNELDIR) M=$(PWD) modules
Endif


static struct miscdevice miscdev= {
        /*
         * We don't care what minor number we end up with, so tell the
         * kernel to just pick one.
         */
        MISC_DYNAMIC_MINOR,
        /*
         * Name ourselves /dev/UniPro.
         */
        "miscdev",
        /*
         * What functions to call when a program performs file
         * operations on the device.
         */
        &fops
};

misc_register(&miscdev);